Browser-based speech studio · 2024
Text2Vox
A browser-based speech studio that transforms text into natural audio with adjustable voice controls, instant preview, and downloadable output.
- Role
- Frontend, audio integration
- Duration
- 2 weeks
Give every sentence a voice.
Text2Vox makes hosted speech models feel like a small, understandable creative tool. It brings writing, voice controls, preview, and export together in a single browser workflow.
The context
Many text-to-speech interfaces either provide too little control or expose technical settings without explaining how they affect the result.
The approach
I designed a clear sequence from text input to voice selection, tone adjustments, preview, and download, with Web Audio handling responsive playback inside the browser.
Architecture
A simplified view of how the main product layers exchange data.
Responsibilities and data flow are shown here without infrastructure noise.
Built with
Core features
- 01
Hosted natural-speech model integration
- 02
Adjustable voice, pitch, pace, and style
- 03
Immediate browser playback
- 04
Downloadable output for offline use
What required care
- Keeping audio feedback immediate across browsers
- Handling inference delay without blocking the interface
- Presenting voice controls in plain language
What the work clarified
- Fast preview loops matter more than an exhaustive control surface.
- Audio tools benefit from visible states and forgiving defaults.
Key product flows
Three direct paths that show how the product helps someone move from intent to outcome.
Writing workspace
A focused product flow designed around one clear user outcome.
Voice controls
A focused product flow designed around one clear user outcome.
Export flow
A focused product flow designed around one clear user outcome.