Research programme
Software that knows when it doesn't know.
EverHeard captions speech and flags meaningful sounds entirely on-device, for people who are Deaf or hard of hearing. Our research programme takes on the problem the product does not yet solve: teaching an on-device system to recognise when it isn't sure, and to say "unclear" or stay silent rather than present a confident error the reader cannot catch.
Why it matters
A hearing person notices a bad caption, because they also heard the words. A Deaf or hard-of-hearing reader has only the caption. That makes a fluent, confident mistake more damaging than a gap — it is trusted precisely because it reads well.
Small speech models that run on a phone make this worse. They offer no dependable measure of their own certainty, and they can produce plausible text from silence or background clatter. The same is true of naming a speaker or identifying a sound: a confident guess, with nothing behind it.
What we are researching
One on-device system, advanced along five aims:
- Calibrated uncertainty and principled abstention across captions, speaker identity and sound events — the core of the work
- Open-set speaker recognition: rejecting the unknown voice instead of forcing a match
- Streaming, persistent identity: who is speaking, live, as a conversation grows
- Robustness in real rooms, with noise, overlap and non-speech
- On-device efficiency, so all of it fits within a phone's compute, thermal and battery budget
Progress is measured by a trust-aware risk–coverage curve, which rewards correct silence and penalises confident errors, rather than by word error rate alone.
Status. R&D with a defined roadmap. Calibrated abstention and reliable speaker identity are research goals, not shipped features. A provisional patent covering the calibrated-abstention method was filed in September 2026 (patent pending).
What we bring
- A shipped, fully on-device system across iPhone, iPad and Mac, with an Apple Watch companion
- On-device engines for speech recognition, sound-event classification and speaker embedding
- The policy layers around them: hallucination filtering, noise gating, honest three-tier sound labelling with a safety override
- Real-device field logs, an ongoing TestFlight beta, and a connection to the Deaf and hard-of-hearing community for validation
- A GPU compute grant from Lambda, through NVIDIA Inception
- The ability to ship what the research produces to real users
What we're looking for
We are seeking a university research partner for an NSF STTR collaboration: a lab to lead the recognition science, with a named senior researcher as day-to-day technical lead, while BTC Media Labs leads the product, the on-device engineering and the deployment.
For a partner lab, this offers co-authored publications on a genuinely unsolved problem, a real deployment platform and field data for evaluation, graduate-student funding through the subaward, and broader impacts grounded in an underserved accessibility need.
We are currently in discussions with research institutions, and a second NSF project pitch is in preparation.
Researching this with us
If your lab works on speaker recognition, diarisation, robust speech, or uncertainty-aware machine learning, we'd like to hear from you. Tell us about your group's relevant work, whether you could commit named effort, your rough availability, and any prior SBIR or STTR experience. We'll send a one-page sketch of the aims and propose a call.
Privacy, by construction
The product processes everything on the device; nothing leaves it in normal use. Any research data is gathered only through an explicit, consented development channel, never from the distributed app. We treat that constraint as part of the research problem rather than an obstacle to it.