Skip to content

Proposal: Indic code-switched data (Hindi-English / Tamil-English) + code-switch metric #177

Description

@balaguhanesh

#175 adds Hindi data, but monolingual. In production Indian-market voice agents, callers code-switch intra-sentence ("mera appointment reschedule karna hai next Tuesday ko") — exactly where ASR degrades hardest. The ServiceNow code-switching benchmark covers ES/FR/DE-English; no Indic pairs exist anywhere.

I'd like to contribute, in stages:

  1. Hindi-English + Tamil-English code-switched utterance sets across existing domains, following Adding Korean and Hindi data #175's conventions (I'm a native speaker of both; I build STT/TTS eval harnesses for code-switching at 2care.ai)
  2. Code-Switch F1 scoring alongside WER — accuracy at switch-point words, so boundary degradation isn't averaged away
  3. (Follow-up PR) a code-switch perturbation for the perturbation pipeline

Before building: is this something you'd take, and is anyone working on it internally? If green-lit, first set + scoring ready for review in ~a week.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions