Showing posts with label conference. Show all posts
Showing posts with label conference. Show all posts

Tuesday, September 23, 2025

Two papers accepted at ICSigSys 2025!

I’m thrilled to announce that two of my recent submissions have been accepted for presentation at ICSigSys 2025. Both pieces push the envelope in speech processing, blending self-supervision, domain adaptation, and cross-lingual storytelling to tackle real-world challenges. Here’s a closer look at each paper and what makes them special.


Semi-Supervised Acoustic Scene Classification with Label Smoothing and Hard Samples Identification

Acoustic scene classification (ASC) remains a cornerstone task in environmental audio understanding. In this work, we introduce a semi-supervised framework that leverages large amounts of unlabeled audio while focusing the model’s attention on the most informative samples.

  • We employ label smoothing to soften the target distribution, reducing overconfidence in noisy or ambiguous audio segments.
  • We design a hard sample identification strategy that dynamically selects challenging clips during training, guiding the model to learn discriminative features more robustly.
  • Our experiments on standard ASC benchmarks demonstrate a consistent performance boost over fully supervised baselines, especially under limited labeled data regimes.

By integrating these elements, our model adapts more gracefully to new acoustic conditions and requires fewer manual annotations. I look forward to sharing detailed analyses of embedding drift, class confusion matrices, and ablation studies during the conference.


Indonesian Folklore Storytelling in Japanese Language with Text-to-Speech

Cross-lingual storytelling opens up rich cultural exchanges, but high-quality narration across languages is still under-explored. This paper presents a text-to-speech (TTS) system that brings Indonesian folklore into Japanese, preserving narrative style and emotional nuance.

  • We start with a Japanese TTS backbone fine-tuned on expressive speech corpora to capture intonation and rhythm.
  • We build a lightweight text conversion pipeline that maps Indonesian story scripts to Japanese text, retaining metaphorical and cultural references.
  • We evaluate the generated speech with both objective metrics (e.g., Mel-cepstral distortion) and subjective listening tests, showing high naturalness and emotional congruence.

This work bridges two rich oral traditions and showcases how speech technologies can make cultural content accessible across language barriers. I’m excited to demo sample audio clips and discuss potential extensions to other language pairs.


Looking Ahead

  • Expand the hard sample identification strategy to multilingual acoustic scenes.
  • Incorporate emotion labels into the semi-supervised ASC framework for more nuanced predictions.
  • Generalize the folklore TTS pipeline to handle low-resource languages with minimal parallel data.

I’m grateful to my co-authors and colleagues at NAIST for their support and valuable discussions. If you’ll be at ICSigSys 2025, please stop by our sessions—we’d love to hear your feedback and explore collaborations.

Stay tuned for preprint links, code releases, and audio demos. Your insights will help shape the next phase of this research journey!

Thursday, August 28, 2025

Two Papers Accepted at APSIPA 2025 in Singapore

I’m thrilled to share that two of my papers have been accepted for presentation at the APSIPA Annual Summit and Conference 2025, taking place in vibrant Singapore this October. This marks a significant milestone for our research team and underscores our ongoing commitment to advancing speech and audio processing, particularly for health applications.

Accepted Papers

Paper ID Title
120 Dementia Prediction From Speech Signal Using Optimized Prosodic Features
241 Comparison of Solicited and Longitudinal Cough Sounds for Tuberculosis Detection


Paper Summaries

Dementia Prediction From Speech Signal Using Optimized Prosodic Features

This study explores how subtle changes in speech prosody—such as pitch, rhythm, and intensity—can serve as early indicators of dementia. By optimizing feature selection and leveraging machine learning classifiers, our approach achieved a classification accuracy that outperforms several baseline models. We believe this work could pave the way for noninvasive, cost-effective screening tools.

Comparison of Solicited and Longitudinal Cough Sounds for Tuberculosis Detection

In this paper, we examine the diagnostic power of cough sound recordings collected under controlled (“solicited”) versus naturalistic (“longitudinal”) conditions. Our analysis demonstrates that longitudinal data, captured through everyday smartphone use, retains enough acoustic signatures to reliably flag tuberculosis. The findings suggest a path toward scalable, remote health monitoring in resource-limited settings.


Acknowledgements

  • My co-authors and lab mates for their relentless dedication to data collection and algorithm development
  • Funding agencies and institutional support that made this research possible
  • All participants who shared their speech and cough recordings, enabling us to push the boundaries of health diagnostics


Next Steps

  1. Prepare camera-ready manuscripts and finalize supplementary materials
  2. Coordinate travel plans and poster backdrops for Singapore
  3. Schedule rehearsals for the oral presentations
  4. Network with fellow APSIPA attendees to explore collaborations in speech-based health analytics

I look forward to sharing our findings with the APSIPA community and gathering feedback that will fuel the next phase of our research. See you in Singapore!

Monday, January 06, 2025

A paper was accepted at ICAIIC 2025!

Alhamdulillah, our paper has been accepted for the conference ICAIIC 2025. This paper discusses the importance of ensemble learning to improve speech classification accuracy. We proposed performance-weighting methods to evaluate with two variants: using weighted and unweighted accuracies.

Our work is highly beneficial to society, as it will help to improve the performance of speech classification. It may also be generalizable to other domains outside of the speech area.

Several aspects can still be developed further, such as incorporating other weighting methods, and implementation in other datasets as well as in other tasks. We invite readers to collaborate in addressing the unresolved questions above.

We extend our gratitude to AIST for their full support of our research, NEDO and JST for research funding.

Happy reading. We welcome your feedback. See you in Fukuoka!

URL for downloading the paper: (will be given after it is available or contact me to get the accepted version).


ICAIIC 2025 paper


Thursday, September 12, 2024

Two papers got accepted in TENCON 2024

Two of my papers were accepted in TENCON 2024. Here is the list of titles:

  1. Multi-label Emotion Share Regression From Speech Using Pre-Trained Self-Supervised Learning Models
  2. Evaluating Hyperparameter Optimization for Machinery Anomalous Sound Detection

The first paper talks about emotion (share) recognition, meaning how to predict more than a single emotion from utterance. It differs from general speech emotion recognition (SER), although we can select n highest probabilities from SER. In the former, the total share should be 1 (or 100%). In the latter, the probabilities of each emotion category are independent, i.e., each could have 0.85 and 0.75 of probabilities. Usually, the highest probability is selected.

Here is a more detailed example.

Emotion (share) recognition

Angry: 0.54

Fear: 0.43

Other: 0.03

Speech emotion recognition

Angry: 0.64

Fear: 0.53

Sad: 0.23

In the first, the sum up of all probabilities is 1; this is not the case for the second approach (SER).

In the second article, I optimized anomalous machine sound detection via Optuna. The results on two different databases show different values of optimal parameters; however, the top three parameters to optimize remain the same (learning rate, patience, and type of loss function).

See you in Singapore, inshallah! 


Thursday, August 22, 2024

A paper was accepted at ACM MM 2024 Workshop

 


I am delighted to show that my paper was accepted at the ACM MM 2024. This was my first ACM paper and was written by myself solely (solo author). Here is an abstract from the screenshot of the paper above.

Abstract

 Automatic social perception recognition is a new task to mimic the measurement of human traits, which was previously done by humans via questionnaires. We evaluated unimodal and multimodal systems to predict agentive and communal traits from the LMU-ELP dataset. We optimized variants of recurrent neural networks from each feature from audio and video data and then fused them to predict the traits. Results on the development set show a consistent trend that multimodal fusion outperforms unimodal systems. The performance-weighted fusion also consistently outperforms mean and maximum fusions. We found two important factors that influence the performance of performance-weighted fusion. These factors are normalization and the number of models.

Once the link to the paper is available in the ACM Library, I will put the link here.

Link: https://dl.acm.org/doi/10.1145/3689062.3689082.

Behind The Scene and Tips!

I have participated in the MuSe challenge (Multimodal Sentiment Analysis Challenge and Workshop) for several years. This challenge, along with other challenges in conferences like ICASSP and Interspeech, usually provides ba aseline program (code) and the respected dataset (e.g., ComParE). From this baseline, we can further analyze, make experiments, and often get new ideas to implement. My idea for that challenge (social perception challenge) is two parts: parameter optimization and multimodal fusion. I implement a lot of ideas (e.g., tuning more than 15 parameters) and some works. Once I get improvement with consistent results/phenomena (science must be consistent!), I documented my work and submitted a paper. This time, my paper got accepted!

See you in Melbourne, inshallah!

Friday, December 22, 2023

Tentang Gift dan Ghost Authorship

If research misconduct occurs, including guest/gift authorship, the integrity of a researchers is questionable; this can PERMANENTLY and NEGATIVELY affect their career.

 

ghost authorsip

Salah salah praktek tercela di bidang akademik adalah ghost dan gift authorship, menuliskan nama mereka yang tidak berkontribusi di dalam penulisan paper. Penerbit seperti Elsevier dan Springer memiliki peraturan ketat dalam kasus ini, sekali praktek ini ditemukan, karya yang bersangkutan bisa  ditarik. Saya sendiri sudah beberapa kali melaporkan kasus ini, tak peduli rekan sejawat, atasan, atau orang lain. Ada kalanya, mungkin, anda "dipaksa" menuliskan nama teman anda, baik kenalan, teman se-lab, se-kampus, atau se-jurusan. Untuk dimasukkan menjadi penulis (authors), setidaknya ada tiga syarat [1]:
1. Kontribusi subtansial dalam riset
2. Ikut menulis draft
3. Menyetujui versi final dari draft

Cara mengetes ghost author, menurut elsevier, adalah bahwa semua penulis mempunyai kemampuan dan kewajiban yang sama dalam mempertahankan ide di dalam tulisannya. Oleh karena itu, jika ada seorang penulis (co-authors) yang tidak bisa menjelaskan atau menjawab pertanyaan terkait tulisannya, bisa jadi penulis itu adalah ghost, guest atau gift author.

Compliance is more than just obeying laws.

Dalam publikasi ilmiah, kita memiliki code of conduct. Dalam setiap submisi artikel, kita diwajibkan menuliskan kontribusi setiap penulis dalam artikel. Secara common sense, setiap mereka yang namanya ada pada paper pasti berkontribusi. Hukum penulisan ini wajib kita patuhi. Bahkan diatasnya, ada etika yang seharusnya kita taati juga: menghindari research misconduct termasuk gift, ghost atau honorary authorship ini.

Jika ada orang lain yang meminta namanya ditulis dalam sebuah paper, sebaliknya ada juga orang yang melarang namanya ditulis dalam sebuah paper. Seorang teman S3 pernah bercerita bahwa professornya, meminta namanya tidak dimasukkan dalam publikasi teman saya tersebut karena dia tidak berkontribusi sama sekali.


Institusi kita punya visi dan misi mulia yang ingin diraih. Kita tidak bisa menghalalkan segala cara untuk meraih visi dan misi itu. You CANNOT choose just any method to achieve your goal. There is a "path we should take" among them. Goal adalah visi institusi kita yang ingin kita capai. Jelas disini bahwa ghost dan gift authorship tidak termasuk dalam "path we should take".
 
Rezeki tidak hanya dari insentif publikasi. Masih banyak jalan dan pintu rezeki berkah dan halal lainnya daripada memasukkan nama istri, teman, atau atasan yang tidak berkontribusi pada publikasi ilmiah kita. Praktek tercela ini membahayakan insitusi kita. Semoga kita terhindar dari praktek tercela ini.

Ghost authorship untuk memperbesar peluang diterimanya paper

Ini disebut efek Chaporone [3]. Penulis Chaperone adalah penulis senior yang telah menulis beberapa journal, katakanlah di jurnal A. Untuk memperbesar kans diterima di journal A, penulis junior mengajak atau memasukkan nama penulis Chaperone saat submisi ke Jurnal A. Tujuannya adalah untuk mempertinggi kans diterima di jurnal A.

Ada kasus dimana peneliti yang sudah meninggal tetap dilibatkan dalam pembuatan paper. Kasus ini [2], ditengarai untuk memperbesar kans diterimanya paper tersebut di sebuah jurnal. Anggaplah peneliti yg sudah meninggal tsb adalah peneliti terkenal, misal peraih nobel. Dengan memasukkan namanya sebagai co-author maka kans diterimanya sebuah paper dalam jurnal mungkin akan bertambah besar. Hal yang tidak mungkin dilakukan oleh peneliti yang sudah meninggal adalah pada persyaratan nomor 3 authorship di atas: menyetujui versi final. Pun demikian, hal ini (memasukkan penulis yang sudah meninggal) bisa saja dilakukan bila penulis yang telah meninggal tersebut benar-benar berkontribusi dan disebutkan dalam "Aknowledgement" bahwa salah satu penulis telah meninggal sebelum paper diterbitkan.

Kenapa ada praktek Ghost Authorship, khususnya di negeri kita?

Seorang sejawat bertanya, kenapa kondisi ideal (tidak ada ghost authorship) tidak bisa diterapkan di, khususnya, negeri kita tercinta. Banyak faktor. Diantara banyak faktor, menurut saya yang paling penting adalah mental peneliti dan kecukupan ekonominya. Di negeri kita, mental peneliti belum terbentuk secara ideal. Alih-alih melakukan "impactful research"; yang dilakukan peneliti adalah bagaimana mendapatkan cuan dari penelitian, entah itu dari insentif, kenaikan pangkat, dll. Disini mental peneliti yang bersangkutan bermasalah. Bisa jadi karena tidak ada pendidikan "Compliance and Researcher Ethics" untuk para peneliti (di institusi saya, setiap peneliti diwajibkan mengambil e-learning ini setiap tahunnya dan wajib lulus ujian e-learning tsb). Kondisi ini diperparah dengan kecukupan ekonomi peneliti yang bersangkutan. Selama kebutuhan dasar belum terpenuhi (sandang, pangan, papan), maka dia akan mencari segala cara (dan mungkin saja menghalalkan segala cara) untuk memenuhi kebutuhan tersebut. Salah satunya dengan meminta rekannya untuk memasukkan namanya saat publikasi. Agar, ketika mendapat insentif, dia juga ikut kecipratan. Sekaligus mempercepat proses kenaikan pangkat.


Lalu, Bagaimana solusinya?

Alih-alih menuliskan teman atau orang lain sebagai ghost authorship, kita bisa menawari mereka untuk berkontribusi, misalnya:

  • Funding, membiayai biaya publikasi
  • Proof-read, misalnya menemukan minimal 10 kesalahan dalam draft dan merevisinya

Dengan cara itu, seorang menjadi layak menjadi co-author. Dengan catatan, sekali lagi, kontribusinya substansial (dua diatas adalah contoh kotribusi yang substansial).

 

Referensi: 

[1] https://www.jscpt.jp/eng/journal/kitei.html

[2] https://retractionwatch.com/2024/02/16/highly-cited-scientist-published-dozens-of-papers-after-his-death/

[3] V. Sekara, P. Deville, S. E. Ahnert, A. L. Barabási, R. Sinatra, and S. Lehmann, “The chaperone effect in scientific publishing,” Proc. Natl. Acad. Sci. U. S. A., vol. 115, no. 50, pp. 12603–12607, 2018, doi: 10.1073/pnas.1800471115.

 

Wednesday, January 19, 2022

Choosing Journals and Conferences for Publication: Google Top 20 (and h5-index > 30)

If you want to publish your academic paper in a conference or journal, you may be confused about to which conference or journal you should submit your papers to. This short article may help you. To be categorized as a "reputable journal", my institution required two indicators below.

  1. It appears in Google Top 20 (all categories, categories, and sub-categories)
  2. It has Google H5-index > 30
For the first reason, it makes sense. The top twenty are the top 20 journals and conferences (mixed) which have the highest h5-index. I don't know the reason for the second reason why my institution chooses 30 as the limit of h5-index for "more incentive". It still makes sense since the higher h5-index means the higher impact.

From those two indicators, I choose the first as the main criteria for selecting publication. Here are five top 20 journals and conferences from all categories, categories, and two sub-categories in my field.


Google Top-20 (all categories)


For choosing categories, click "VIEW ALL"
https://scholar.google.com > Metrics > VIEW ALL.

Google Top 20 Category Engineering and Computer Sciences



Google Top 20 Category Life Sciences and Earth Sciences



Google Top 20 Sub-category: Signal Processing



Google Top 20 Sub-category: Acoustic and Audio



This guide for selecting criteria is not mandatory in my constitution. But they will give more bonus to the researchers if their publications are ranked by one or both criteria above (more bonus for both, maybe).

Wednesday, August 05, 2015

Signal Enhancement By Single Channel Source Separation

Most gadgets and electronics devices are commonly equipped with single microphone only. This is difficult task in source separation world which traditionally required more sensors than sources to achieve better performance. In this paper we evaluated single channel source separation to enhance target signal from inteferred noise. The method we used is non-negative matrix factorization (NMF) that decompose signal into its components and find the matched signal to target speaker. As objective evaluation, coherence score is used to measure the perceptual similarity from enhanced to original one. It show the extracted has 0.5 of average coherence that shows medium correlation between both signals.

The following slides talk a bit about signal Signal enhancement by single channel source separation principle. You can grab the full paper here.


Friday, July 03, 2015

Extracting Sound From Multiple Sources in Anechoic Room


Ini adalah revisi dari poster saya sebelumnya. What I learned? Apa yang saya perbaiki?
  • Judul harus menarik dan spesifik, tidak lebih dari 10 kata
  • Tambahkan ABSTRAK
  • Hindari Jargon (kata-kata spesifik) dan akronim
  • Keep it Simple!
  • Tunjukkan the BIG idea, the big deal!
  • Tunjukkan apa yang menarik dan kenapa
  • Next Step, apa?
  • Font minimal 28, idealnya 36 
  • Perhatikan flow (kiri ke kanan) dan balance layout 
  • Keywords! orang menemukan poster/paper kita dengan bantuan keyword (termasuk anda ketika menemukan tulisan ini)
 Have a good poster and snapshot!