OpenAI’s TTS API Now Takes Tone Instructions

Every voice agent built on generic TTS has been bleeding users the moment it sounds like a robot reading terms and conditions.

Flat, robotic voice output has been killing conversion for AI agents

Developers building voice agents have had one option for tone: whatever the model felt like. Swapping delivery style meant re-recording, hiring voice actors, or shipping a product that felt cold at exactly the wrong moment.

You now pass a tone instruction the same way you pass a prompt

Introducing next lets developers pass a plain-language instruction alongside the text input, something like “speak like a sympathetic customer service agent,” and the model adjusts cadence, warmth, and delivery accordingly. You call the API, include your tone directive in the request, and receive audio output that matches the emotional register you specified. The input is text plus instruction; the output is audio that no longer sounds like every other voice agent on the market.

Voice product teams feel this the most

  • Conversational AI developers who lose users in the first ten seconds because the agent sounds indifferent to the actual problem being described
  • Customer support teams deploying voice bots where tone mismatch between a tense caller and a flat response creates escalations instead of resolutions
  • Consumer app builders who need a character-consistent voice across onboarding, error states, and celebration moments without maintaining separate audio assets

ElevenLabs has held significant ground in expressive TTS precisely because OpenAI’s API offered no tone control. That gap just closed, and every voice product currently routing audio through a third-party provider has a reason to re-evaluate its stack this quarter.

What you can do with tone-directed TTS today

  • Direct the model to sound urgent for time-sensitive alert systems
  • Set a calm, measured register for mental health or wellness applications
  • Build persona-consistent voice layers across multiple product touchpoints
  • Test tone variants against user retention data without re-recording a single line

Pricing is usage-based through the OpenAI API — check the aineedthat.com directory for current rate details.

Tone instruction quality degrades on long-form output, so complex emotional arcs across multi-minute audio are not yet reliable.

ElevenLabs remains the stronger choice for fine-grained voice cloning. For teams already inside the OpenAI ecosystem, the switching cost to maintain a separate TTS provider just became harder to justify.

The gap between expressive and generic voice AI just collapsed

This is one of the more quietly significant API updates of the year for anyone building voice-first products. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.