Selected system
Low-resource speech recognition, built from the data layer upward.
Turkmen ASR — Open Data & Evaluation Stack
A public, reproducible Turkmen ASR stack that confronts an overlooked evaluation problem: machine-generated references can hide model error. The work connects dataset curation, human evaluation, transcript correction, LoRA fine-tuning, and transparent error analysis.
- Fine-tuned WER
- 24.44%fixed 300-clip held-out evaluation
- Zero-shot baseline
- 124.61%same clips, decoder, and normalization
- Relative WER reduction
- 80%vs. zero-shot Whisper Large v3 Turbo
- Training corpus
- 251 hnatural spoken Turkmen with LoRA fine-tuning
