Showing posts with label academic paper. Show all posts
Showing posts with label academic paper. Show all posts

Tuesday, September 23, 2025

Two papers accepted at ICSigSys 2025!

I’m thrilled to announce that two of my recent submissions have been accepted for presentation at ICSigSys 2025. Both pieces push the envelope in speech processing, blending self-supervision, domain adaptation, and cross-lingual storytelling to tackle real-world challenges. Here’s a closer look at each paper and what makes them special.


Semi-Supervised Acoustic Scene Classification with Label Smoothing and Hard Samples Identification

Acoustic scene classification (ASC) remains a cornerstone task in environmental audio understanding. In this work, we introduce a semi-supervised framework that leverages large amounts of unlabeled audio while focusing the model’s attention on the most informative samples.

  • We employ label smoothing to soften the target distribution, reducing overconfidence in noisy or ambiguous audio segments.
  • We design a hard sample identification strategy that dynamically selects challenging clips during training, guiding the model to learn discriminative features more robustly.
  • Our experiments on standard ASC benchmarks demonstrate a consistent performance boost over fully supervised baselines, especially under limited labeled data regimes.

By integrating these elements, our model adapts more gracefully to new acoustic conditions and requires fewer manual annotations. I look forward to sharing detailed analyses of embedding drift, class confusion matrices, and ablation studies during the conference.


Indonesian Folklore Storytelling in Japanese Language with Text-to-Speech

Cross-lingual storytelling opens up rich cultural exchanges, but high-quality narration across languages is still under-explored. This paper presents a text-to-speech (TTS) system that brings Indonesian folklore into Japanese, preserving narrative style and emotional nuance.

  • We start with a Japanese TTS backbone fine-tuned on expressive speech corpora to capture intonation and rhythm.
  • We build a lightweight text conversion pipeline that maps Indonesian story scripts to Japanese text, retaining metaphorical and cultural references.
  • We evaluate the generated speech with both objective metrics (e.g., Mel-cepstral distortion) and subjective listening tests, showing high naturalness and emotional congruence.

This work bridges two rich oral traditions and showcases how speech technologies can make cultural content accessible across language barriers. I’m excited to demo sample audio clips and discuss potential extensions to other language pairs.


Looking Ahead

  • Expand the hard sample identification strategy to multilingual acoustic scenes.
  • Incorporate emotion labels into the semi-supervised ASC framework for more nuanced predictions.
  • Generalize the folklore TTS pipeline to handle low-resource languages with minimal parallel data.

I’m grateful to my co-authors and colleagues at NAIST for their support and valuable discussions. If you’ll be at ICSigSys 2025, please stop by our sessions—we’d love to hear your feedback and explore collaborations.

Stay tuned for preprint links, code releases, and audio demos. Your insights will help shape the next phase of this research journey!

Thursday, August 28, 2025

Two Papers Accepted at APSIPA 2025 in Singapore

I’m thrilled to share that two of my papers have been accepted for presentation at the APSIPA Annual Summit and Conference 2025, taking place in vibrant Singapore this October. This marks a significant milestone for our research team and underscores our ongoing commitment to advancing speech and audio processing, particularly for health applications.

Accepted Papers

Paper ID Title
120 Dementia Prediction From Speech Signal Using Optimized Prosodic Features
241 Comparison of Solicited and Longitudinal Cough Sounds for Tuberculosis Detection


Paper Summaries

Dementia Prediction From Speech Signal Using Optimized Prosodic Features

This study explores how subtle changes in speech prosody—such as pitch, rhythm, and intensity—can serve as early indicators of dementia. By optimizing feature selection and leveraging machine learning classifiers, our approach achieved a classification accuracy that outperforms several baseline models. We believe this work could pave the way for noninvasive, cost-effective screening tools.

Comparison of Solicited and Longitudinal Cough Sounds for Tuberculosis Detection

In this paper, we examine the diagnostic power of cough sound recordings collected under controlled (“solicited”) versus naturalistic (“longitudinal”) conditions. Our analysis demonstrates that longitudinal data, captured through everyday smartphone use, retains enough acoustic signatures to reliably flag tuberculosis. The findings suggest a path toward scalable, remote health monitoring in resource-limited settings.


Acknowledgements

  • My co-authors and lab mates for their relentless dedication to data collection and algorithm development
  • Funding agencies and institutional support that made this research possible
  • All participants who shared their speech and cough recordings, enabling us to push the boundaries of health diagnostics


Next Steps

  1. Prepare camera-ready manuscripts and finalize supplementary materials
  2. Coordinate travel plans and poster backdrops for Singapore
  3. Schedule rehearsals for the oral presentations
  4. Network with fellow APSIPA attendees to explore collaborations in speech-based health analytics

I look forward to sharing our findings with the APSIPA community and gathering feedback that will fuel the next phase of our research. See you in Singapore!

Monday, August 04, 2025

Perbedaan luaran institusi pendidikan dan institusi riset

Empat bulan ini saya menjalani profesi baru, menjadi asisten profesor di sebuah universitas. Sebelumnya saya bekerja di institusi riset. Meski perubahan profesi ini terlihat smooth, ada hal dasar yang membedakan antar keduanya.

Persamaan luaran institusi riset dan pendidikan terletak pada publikasi. Dua-duanya menilai publikasi sebagai luaran. Namun ada perbedaan mendasar antar keduanya, yakni pada sisi authorship. Sebagai periset, menjadi penulis utama adalah tolok ukur keberhasilan periset yang bersangkutan. Hal ini berbeda dengan profesi dosen di institusi pendidikan.

Di institusi pendidikan, tolok ukur utama keberhasilan adalah ketika bisa mendidik mahasiswa untuk menjadi penulis pertama  dalam sebuah publikasi, entah itu conference paper atau pun jurnal. Seperti dalam tulisan saya sebelumnya, menjadi penulis pertama artinya menjadi kontributor utama. Disini lah letak keberhasilan itu, mendidik mahasiswa untuk berkontribusi dalam sains dan riset. Bukan menjadikan dirinya sebagai penulis pertama. 
 

Bagaimana seharusnya menilai kinerja periset dan dosen?


Karena tolok ukur keberhasilan yang diusulkan di atas tadi berdasarkan keberhasilan menjadikan mahasiswa sebagai penulis utama, maka setidaknya ada dua kriteria untuk menilai periset dosen. Kriteria pertama dengan banyaknya (kuantitas), kriteria kedua dengan seberapa baik kualitas paper atau journal yang diterbitkan. Untuk kriteria pertama cukup jelas, misalnya berapa jurnal publikasi per tahun. Untuk kriteria ada beberapa metrik/standard yang bisa dipakai. Cara pertama, yang saya usulkan, adalah dengan memakai metrik Google Scholar, yakni kualitas publikasi dinilai dari masuk tidaknya jurnal atau conference (keduanya dipukul rata) tersebut dalam Google Top 20. Misalnya bidang saya, acoustic and sound, ada pada list berikut. Cara ini cukup efisien dan fair, tidak peduli entah dia jurnal atau conference paper. Cara kedua yakni dengan menggunakan quartile Scopus, yakni publikasi (hanya jurnal) harus masuk antara Q4 sampai Q1. Semakin tinggi nilai Q-nya, semakin tinggi bobot kualitasnya. Namun ada kelemahan dalam cara kedua ini, yakni bagaimana menilai conference paper? Padahal conference paper dewasa ini keterbaruannya lebih tinggi daripada jurnal karena frekuensi penyelenggarannya tahunan.


Masalah selanjutnya adalah dalam penilaian kinerja periset atau dosen. Karena perbedaan authorship luaran diatas, maka penilaiannya harus dibedakan pula. Tidak seharusnya dosen dinilai dari sisi penulis pertama; sebaliknya hal tersebut berlaku pada periset. Dosen hendaknya dinilai dari kualitas dan kuantitas publikasi yang dihasilkan, yang besar kemungkinan berbanding lurus dengan jumlah mahasiswa yang diluluskannya.

Tuesday, January 14, 2025

A paper was accepted at 2025 ICASSP Workshop!

Alhamdulillah, our paper has been accepted for a satellite workshop of ICASSP 2025. This paper discusses "Pathological Voice Detection From Sustained Vowels: Handcrafted vs. Self-supervised Learning". We proposed to examine pathological voice detection from sustained vowels (/a/, /i/, /u/) both using acoustic features and self-supervised learning (SSL) models. We also evaluated early fusion (feature concatenation) and decision-level ensemble learning for both types of features.

Our work is highly beneficial to society, as it will help to improve the performance of pathological voice detection. 

Several aspects were evaluated in this research project: evaluation of different vowels (which one leads to better results), evaluation of different acoustic and SSL features, and ensemble learning results.

Future work could tackle the limitations of the F1 score AUC by using more recent metrics like the Matthew correlation coefficient (MCC), which considers true and false positives and negatives.

Since the nature of the problem of detecting pathological voices can be classified as anomaly detection, future work can also be accomplished to observe the effectiveness of anomaly detection methods for pathological voice detection.

We extend our gratitude to AIST for their full support of our research, and to NEDO and JST for research funding.

Happy reading. We welcome your feedback. See you in Hyderabad!

2025 ICASSP Workshop Paper


URL for downloading the paper: (will be given after it is available or contact me to get the accepted version). 

Project repository: https://github.com/bagustris/svd-exploration

Monday, January 06, 2025

A paper was accepted at ICAIIC 2025!

Alhamdulillah, our paper has been accepted for the conference ICAIIC 2025. This paper discusses the importance of ensemble learning to improve speech classification accuracy. We proposed performance-weighting methods to evaluate with two variants: using weighted and unweighted accuracies.

Our work is highly beneficial to society, as it will help to improve the performance of speech classification. It may also be generalizable to other domains outside of the speech area.

Several aspects can still be developed further, such as incorporating other weighting methods, and implementation in other datasets as well as in other tasks. We invite readers to collaborate in addressing the unresolved questions above.

We extend our gratitude to AIST for their full support of our research, NEDO and JST for research funding.

Happy reading. We welcome your feedback. See you in Fukuoka!

URL for downloading the paper: (will be given after it is available or contact me to get the accepted version).


ICAIIC 2025 paper


Thursday, September 12, 2024

Two papers got accepted in TENCON 2024

Two of my papers were accepted in TENCON 2024. Here is the list of titles:

  1. Multi-label Emotion Share Regression From Speech Using Pre-Trained Self-Supervised Learning Models
  2. Evaluating Hyperparameter Optimization for Machinery Anomalous Sound Detection

The first paper talks about emotion (share) recognition, meaning how to predict more than a single emotion from utterance. It differs from general speech emotion recognition (SER), although we can select n highest probabilities from SER. In the former, the total share should be 1 (or 100%). In the latter, the probabilities of each emotion category are independent, i.e., each could have 0.85 and 0.75 of probabilities. Usually, the highest probability is selected.

Here is a more detailed example.

Emotion (share) recognition

Angry: 0.54

Fear: 0.43

Other: 0.03

Speech emotion recognition

Angry: 0.64

Fear: 0.53

Sad: 0.23

In the first, the sum up of all probabilities is 1; this is not the case for the second approach (SER).

In the second article, I optimized anomalous machine sound detection via Optuna. The results on two different databases show different values of optimal parameters; however, the top three parameters to optimize remain the same (learning rate, patience, and type of loss function).

See you in Singapore, inshallah! 


Thursday, August 22, 2024

A paper was accepted at ACM MM 2024 Workshop

 


I am delighted to show that my paper was accepted at the ACM MM 2024. This was my first ACM paper and was written by myself solely (solo author). Here is an abstract from the screenshot of the paper above.

Abstract

 Automatic social perception recognition is a new task to mimic the measurement of human traits, which was previously done by humans via questionnaires. We evaluated unimodal and multimodal systems to predict agentive and communal traits from the LMU-ELP dataset. We optimized variants of recurrent neural networks from each feature from audio and video data and then fused them to predict the traits. Results on the development set show a consistent trend that multimodal fusion outperforms unimodal systems. The performance-weighted fusion also consistently outperforms mean and maximum fusions. We found two important factors that influence the performance of performance-weighted fusion. These factors are normalization and the number of models.

Once the link to the paper is available in the ACM Library, I will put the link here.

Link: https://dl.acm.org/doi/10.1145/3689062.3689082.

Behind The Scene and Tips!

I have participated in the MuSe challenge (Multimodal Sentiment Analysis Challenge and Workshop) for several years. This challenge, along with other challenges in conferences like ICASSP and Interspeech, usually provides ba aseline program (code) and the respected dataset (e.g., ComParE). From this baseline, we can further analyze, make experiments, and often get new ideas to implement. My idea for that challenge (social perception challenge) is two parts: parameter optimization and multimodal fusion. I implement a lot of ideas (e.g., tuning more than 15 parameters) and some works. Once I get improvement with consistent results/phenomena (science must be consistent!), I documented my work and submitted a paper. This time, my paper got accepted!

See you in Melbourne, inshallah!

Friday, March 31, 2023

Research Plan 2023

Seperti biasa, bulan April adalah tahun barunya jepang (Fiscal year、年度). Satu tahun (akademik, fiskal, kerja, dll. kecuali pajak) dihitung per 1 April dan berakhir 31 Maret. Jadi kalau bisa, pas hidup di Jepang, lahir tanggal 1 April, meninggal tanggal 31 Maret. Kenapa begitu? Kalau lahir 20 April, anda tidak bisa mengikuti sekolah (SD, SMP, SMA) tahun itu juga kalau belum umurnya. Misalnya umur minimal SD adalah 6 tahun per 1 April. Kalau lahir 20 April berarti nunggu tahun depannya. Begitu juga beasiswa, kerja, dll, hampir semua dimulai per 1 April.

Singkat saja, berikut adalah rencana penelitian saya satu tahun mendatang. Tahun ini saya akan fokus mengembangkan sistem multilingual speech emotion recognition (SER) dengan riset tambahan tentang anomaly detection (vibration and sound) dan melanjutkan riset sebelumnya tentang text-independent speech emotion recognition. Semoga semua luaran tahun ini bisa tercapai, alih-alih bisa melebihi luaran yang diharapkan.




Monday, March 28, 2022

Tools I used for (accelerating) science and research 2022

I made seven publications this year (April 2021 - March 2022) and eight publications last year (April 2020 - March 2021). Here is the list of the tools (and their description) for accelerating my research.

The list is ordered based on its importance.

1. OS: Ubuntu (the most useful part: stfp inside nautilus)

2. vscode (both python and latex, also remote editing)

3. ssh with byoubu (open multiple sessions of the terminal in one instance)

4. Libreoffice (draw, calc, impress, writer)

5. Inkscape (for editing plots)

6. Simplenote (to make notes)

7. Grammarly (paid, mostly inside vscode)

8. Gnumeric (spreadsheet-like)

9. Insync (paid, to host one drive, google drive, dropbox, etc)

10. pdfcrop (to crop your image, automatically!!)

11. clipit (clipboard manager )

12. Mendeley (reference manager, automatic BIB file creation)