# ORF

> Vachan Samiksha · Oral Reading Fluency

- **Role:** Associate ML Scientist I
- **Company:** Wadhwani AI · Education
- **Period:** 2024 – 2025
- **Partners:** Government of Gujarat, Bill & Melinda Gates Foundation

Teachers across large classes need more than a stopwatch and a single score. I contributed to the analysis layer on **Vachan Samiksha (ORF)**, turning child reading audio into **word-level errors**, reading rate, and where a student needs help, as part of a statewide program in **Gujarat government schools**.

## Problem

Under **NIPUN Bharat**, teachers must assess oral reading fluency across large classes, often with a stopwatch and manual word counts. That misses *why* a child struggles: phoneme slips, substitutions, skipped text, or slow but accurate decoding. **Vachan Samiksha (ORF)** replaces the manual workflow with ASR on a student's reading audio, aligned to the passage they were given.

**Take home:** Partners: **Government of Gujarat** (Sarva Shiksha Abhiyan, Vidya Samiksha Kendra) · **Bill & Melinda Gates Foundation** · deployed via **G-Shala** and **ConveGenius** in Gujarat, later **Rajasthan**.

![Oral Reading Fluency program, Wadhwani AI](https://yashwardhan.space/assets/images/work/orf/orf-hero.webp)

*source: Wadhwani AI*

## What I did

### Lexical–sublexical analysis

- Built alignment and scoring experiments on ASR transcripts, separating **whole-word errors** from **phoneme-level decoding mistakes** in Gujarati reading audio.
- Measured how substitutions, hesitations, and statistically odd phoneme sequences show up in classroom recordings.

### Fluency metrics

- Evaluated **character error rate**, **word error rate**, reading duration, and **words-correct-per-minute** as proxies for the program's fluency rubric.
- Worked with the team to connect raw ASR output to teacher-facing signals, not just a single score.

### Gujarati ASR for child speech

- Fine-tuned and benchmarked **Wav2Vec2-style models** on custom Gujarati child-speech data from educational reading scenarios.
- Stress-tested pipelines on noisy field audio, the conditions ORF runs in across Gujarat classrooms.

## About ORF

**Vachan Samiksha (ORF)** is a statewide Wadhwani AI program delivered by multiple teams over several years. The figures below are **program-wide totals** for the full deployment. Per [Wadhwani AI program data (Jan 2026)](https://www.wadhwaniai.org/programs/oral-reading-fluency/), it runs across **Grades 2–8 government schools** in **Gujarat and Rajasthan**. I contributed experiments on lexical–sublexical analysis and Gujarati child-speech ASR on field reading recordings.

- **15M+** program assessments
- **7.9M+** students reached
- **270k+** teachers trained

## Team

- Makarand Tapaswi (Principal ML Scientist · Reporting manager)
- Ayush Deva (Associate ML Scientist II · ORF / Vachan Samiksha lead)
- Vivek Pandey (Associate ML Scientist I · ML team)

## Outcomes

- Built **lexical–sublexical ASR analysis** experiments on Gujarati child-speech reading audio
- Benchmarked **Wav2Vec2-style models** on field recordings for fluency scoring (CER, WER, WCPM)

## Sources

- [Wadhwani AI, Oral Reading Fluency](https://www.wadhwaniai.org/programs/oral-reading-fluency/)
- [Vaachan Samiksha, program overview](https://www.wadhwaniai.org/vaachan-samiksha-leveraging-ai-to-bridge-the-literacy-divide/)

---

[Back to portfolio](https://yashwardhan.space/) · [HTML case study](https://yashwardhan.space/work.html?p=orf)
