Registration closes in 4 d 10 h 44 minAfter that, subscriptions and Stars can't be bought — only renewed by those already inside. 30% off all plans right now.Join now

Practical guide · August 6, 2026

Voice auto-reply in Telegram: transcribing voice notes and answering aloud

Voice auto-reply in Telegram: transcribing voice notes and answering aloud

Plenty of people write in Telegram by voice: it is faster and easier, especially when driving or on the move. If the assistant only understands text, those messages sit unanswered until you listen to them yourself. Here is how HeadHack transcribes incoming voice notes, when it answers with voice, how to pick a voice and what drives the cost.

Short answer. Open @Head_Hack_bot, go to “Profile → Voice replies” and switch the voice side on. Incoming voice notes start being transcribed, and by default the reply comes back in the same format the other person used: voice for voice, text for text. The voice is picked from a list, and previewing is free.

This is part of the guide “A free Telegram auto-reply: how to connect it without Premium”. This page is only about voice.

Why a bot needs a voice

Voice messages have long been an ordinary way to talk: people send them when typing is awkward, when a thought runs longer than a couple of lines, or simply when it is quicker to say it than to write it.

For an auto-reply that creates a gap. An assistant that only understands text does not answer a voice note at all — it waits until you are free and play it back by hand. The whole point of an auto-reply is lost: the other person sees silence where they expected a quick answer.

The voice side closes that gap from both ends: an incoming note becomes text and is handled like any other message, and the reply can come back as voice — in the same format the person used.

How it works

A single voice message travels like this:

  1. the other person sends a voice note in Telegram;
  2. the assistant downloads the recording and transcribes it;
  3. the transcript goes into the ordinary pipeline — the same profiles, prompts, FAQ and documents as for text;
  4. the assistant writes a reply;
  5. the reply goes out as text or is spoken in the chosen voice — whichever you set.

What matters is that transcription does not run down a separate branch of logic. Once recognised, a voice message is no different from a typed one: the same profiles, the same system prompt and the same knowledge base. There is no separate behaviour to configure for voice.

Transcribing incoming voice notes

Recognition has its own switch. Leave it off and voice messages are ignored exactly as they were before the feature existed.

Transcription is not billed in Stars — it is part of the assistant's work. Only speaking the reply can cost anything, and not always: see below.

A limit on length. A voice note is transcribed if it is no longer than a minute. If somebody sends a longer recording the assistant does not go quiet: it politely asks for something shorter, or for text.

Three reply modes

The reply format is chosen once and applies to every conversation:

ModeWhen it fits
Always textVoice notes are transcribed, but the reply always comes as text. Good when speed and the ability to re-read matter
Match the senderVoice for voice, text for text. The default: the conversation stays in the format the other person is used to
Always voiceEvery reply is spoken, even if the question was typed. Good when voice is part of how you talk

“Match the sender” is the default for a reason: it does not push voice on people who would rather read, and does not make people type when they are used to speaking.

Choosing a voice

Voices are shown by ordinary names and a sense of delivery rather than technical codes. The list has male and female voices with different characters — from calm and businesslike to lively and energetic.

They fall into two groups:

  • included in your plan — good synthesis, enough for most jobs;
  • premium (marked with a star) — the most natural intonation, closest to live speech.
Previewing is free. The Play button works for any voice, premium included, and charges nothing. You can work through the whole list and pick the one that suits your style before paying for anything.

Speed, pitch and volume

Three sliders adjust how the chosen voice sounds:

  • speech speed — from noticeably slow to fast;
  • pitch — above or below the base;
  • volume — quieter or louder.

Premium voices have no pitch control — they are synthesised with an intonation of their own, and the height cannot be changed. In that case the pitch slider is simply not shown, while speed and volume work as usual.

After any change, press Play — the sample is re-synthesised with the new settings and you hear the result straight away.

Reply length and what it does to the price

A separate slider sets the maximum reply length — from 100 to 900 characters. It caps both the text reply and the audio: one lever over both how talkative the assistant is and what the voice costs.

It is a ceiling, not a fixed length. The value on the slider means “no longer than”, not “exactly this much”. In practice a reply is nearly always shorter, and you are charged for the characters actually spoken — not for the limit you set.

Which gives a simple rule: the shorter the replies, the cheaper the audio. A typical reply from the assistant comes to around 300 characters and costs noticeably less than the 900-character ceiling. The price per reply recalculates as you drag the slider, so you can find the balance between detail and spend by eye, without sending anything.

If a reply does hit the limit, it is trimmed at a sentence boundary rather than mid-word.

What the plan includes

Every plan includes a set number of voice replies a month. Until that runs out, speaking with the included voices is free.

PlanMessages a monthOf those, by voice
Free5010
Go500100
Plus3,000600
Pro15,0003,000

The remainder is visible right in the settings: how many spoken replies are left from the monthly allowance. The counter resets every month.

Paying beyond the plan with Stars

Anything the plan does not cover is paid for with Telegram Stars from a prepaid balance. That happens in two cases:

  1. an included voice after the monthly allowance runs out — a very small amount per reply;
  2. a premium voice — always, and those replies never spend the monthly allowance.

The balance is topped up in bundles right in the mini app, through ordinary Star payments. The order of charges and the current balance are shown there in their own block: how many Stars you have, how many replies that covers in the chosen voice, and what a typical and a maximum reply will cost.

What happens if the Stars run out. The reply is not lost. If there is nothing to pay for a premium voice with, the reply is spoken in an included voice and you get a notification that the balance is empty. If both the monthly allowance and the balance are exhausted, the reply arrives as text — the other person gets an answer either way.

Setting it up in five steps

  1. Open @Head_Hack_bot and launch the mini app.
  2. Go to “Profile → Voice replies”.
  3. Switch the voice side on and leave incoming transcription enabled.
  4. Pick a reply mode and a voice — listen to a few with the Play button, it is free.
  5. Set the maximum reply length and save.

After that, send the bot a voice message yourself and check that the transcript is right and the reply sounds the way you expected. This is worth adding to the general pre-launch checklist.

Limitations

  • An incoming voice note is recognised if it is under a minute.
  • Recognition and synthesis are tuned for Russian.
  • The owner's voice is not cloned: the ready neural voices from the list are used instead. That is a deliberate choice — cloning requires separate consent and an identity check.
  • Premium voices have no pitch control — only speed and volume.

Common questions

Do I need Telegram Premium?

No. The voice features work like the rest of the assistant — without Premium, without your own bot through BotFather and without a separate server.

Do I pay for transcribing incoming notes?

No. Recognising incoming voice notes is part of the assistant's work. Stars are only charged for speaking replies — and only beyond the monthly allowance or for premium voices.

Can I hear a voice before paying?

Yes, previewing is free for any voice, premium included. You can work through the whole list, change speed and volume and only then decide.

Will the bot speak in my voice?

No. Voice cloning is not used — only the ready voices from the list. You can pick one that suits your style and tune how it sounds.

How do I make the audio cheaper?

Lower the maximum reply length: you are charged for the characters actually spoken, so short replies cost less. The voices included in your plan are several times cheaper than premium.

What happens if someone sends a long voice note?

If the recording is longer than a minute the assistant does not disappear: it asks for something shorter, or for text.

In short

The voice side closes the gap between how people like to write and what the assistant can handle. Incoming voice notes are transcribed free, the reply comes back by default in the format you were addressed in, and the cost of speaking is governed by two understandable levers: the voice you pick and the maximum reply length.

Open @Head_Hack_bot, switch voice replies on, listen to a couple of voices and send the bot a voice message — the whole thing takes a few minutes.