Hi, I’m Kai and I built DeafBench to test what ASR captions actually get wrong #204974
488315
started this conversation in
A Welcome to GitHub
Replies: 1 comment
|
Hou use full is github for website |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Introduce yourself
Hi everyone, I’m Kai and I built DeafBench because WER by itself does not tell me if captions kept the information that actually matters, like a time, dosage, username, Wi-Fi name, confirmation code, speaker, or sound event. I am working on making it a real and reproducible accessibility benchmark instead of another model demo, and I want the results to stay honest, so DeafBench keeps normal ASR scores separate from accessibility-critical recall and I do not claim that it beat a leaderboard unless the official evaluation verifies it.
If anyone here works on ASR, accessibility, benchmarking, Python, or open-source evaluation, I would appreciate someone reviewing the repository and telling me what is missing, what is confusing, or where the evidence needs to be stronger.
Links (optional)
https://github.com/488315/DeafBench
https://488315.github.io/products/deafbench/
Where are you in your GitHub journey?
Making an change with the deaf community
And where are you going next on GitHub?
My next step is improving the public demonstration, adding more authorized real-world accessibility scenarios, and getting independent reviewers to check the scoring, benchmark integrity, and reproducibility before I make bigger claims about the results.
What technical skills or projects are you working on?
I am working on Python ASR evaluation, typed critical-information scoring, acoustic stress testing, frozen benchmark manifests, model adapters, and zero-custody local audit workflows. DeafBench measures ordinary WER, but it also tracks whether models preserve accessibility-critical information, recognize sound events, attribute speakers correctly, and maintain usable latency.
Got a question for us? (optional)
Would anyone with experience in ASR, speech datasets, accessibility, or benchmark design be willing to review DeafBench? I am especially looking for feedback on the scoring, corpus integrity, model comparison lanes, and reproducibility, and I would rather get direct criticism now than make claims the evidence does not support.
All reactions