The problem
A phone call is the last piece of critical infrastructure that still assumes you can hear. Doctors' offices, landlords, employers and benefits offices all say "just give us a call." Relay services exist, but they route a stranger into a private conversation and mostly keep business hours.
Switchboard answers a normal phone number. The hearing caller speaks as usual. The Deaf or hard-of-hearing user reads live captions in a browser and types back; Switchboard speaks their words into the call in a real voice. No third human, no schedule.
One pipeline, three transforms
caller speaks → Vonage ASR → transform → Vonage TTS → caller hears
Every mode is the same pipeline with a different transform in the middle. Each transform is one file in src/modes/:
| Mode | What the transform does | Backing |
|---|---|---|
| Live captions | Passes the caller's words to the browser; speaks the operator's typed reply | None needed |
| Interpreter | Translates each turn and flips the listening language | Claude |
| Navigator | Talks the caller through a phone tree | Claude |
| Call screening | Answers for the owner and asks who is calling and why | Claude |
Adding a mode is adding a file. The keypad menu and the pinned-mode path share one startMode(), so a call that picked "2" and a line pinned to interpreter reach an identical state.
Uncertainty is shown, not hidden
A hearing person says "sorry, what?" a dozen times a day. Captioning systems tend not to: they print their best guess in the same confident type as everything else, and the reader has no way to know which words are solid.
Switchboard's caption mode classifies every utterance before it reaches the screen:
- Below 0.4 confidence — nothing is shown; silence prompts the caller to repeat.
- Below 0.75, or best and runner-up within 0.1 of each other — the caption is marked and the alternatives are printed beside it, with their confidences.
- Otherwise — presented as fact.
The "close call" line in the demo (Tuesday at ten vs Thursday at ten) is the case this exists for.
Built on
- One Cloudflare Worker and one Durable Object. The Express prototype was ported to a singleton DO reached by
idFromName('switchboard'). Singleton, not per-call, because the browser opens the operator view before it knows which call is live, and broadcast fans out to every browser. Sessions snapshot to DO storage every turn and hydrate on wake, so an eviction between webhooks cannot hang up the call. - Vonage Voice API —
input(speech)for recognition,talkfor greetings, and the live-call/talkREST endpoint to inject the operator's typed replies. NCCO callback URLs derive from the origin the webhook arrived on, so there is noPUBLIC_URLto keep in sync with a tunnel. - A hand-rolled RS256 JWT in WebCrypto instead of
@vonage/server-sdk, which signs throughnode:cryptoand has nowhere to run at the edge. - A WebXR presence view (
/ar.html) that places the caller's captions in the room, on the same socket as the console.
The public demo
The instance linked from this page runs the whole pipeline with the telephone swapped for buttons: the staged caller lines carry the confidence and alternatives the recogniser would have returned, and the operator's typed reply takes the exact path it would take into a live call. Caption mode is fully live. The Claude-backed modes show their holding lines rather than run an unauthenticated language model on the public internet.
Built for the DIALED IN Builder Challenge (CreateHER Fest × Vonage), August 2026.