Practical guide · August 6, 2026
Voice auto-reply in Telegram: transcribing voice notes and answering aloud

Plenty of people write in Telegram by voice: it is faster and easier, especially when driving or on the move. If the assistant only understands text, those messages sit unanswered until you listen to them yourself. Here is how HeadHack transcribes incoming voice notes, when it answers with voice, how to pick a voice and what drives the cost.
This is part of the guide “A free Telegram auto-reply: how to connect it without Premium”. This page is only about voice.
Why a bot needs a voice
Voice messages have long been an ordinary way to talk: people send them when typing is awkward, when a thought runs longer than a couple of lines, or simply when it is quicker to say it than to write it.
For an auto-reply that creates a gap. An assistant that only understands text does not answer a voice note at all — it waits until you are free and play it back by hand. The whole point of an auto-reply is lost: the other person sees silence where they expected a quick answer.
The voice side closes that gap from both ends: an incoming note becomes text and is handled like any other message, and the reply can come back as voice — in the same format the person used.
How it works
A single voice message travels like this:
- the other person sends a voice note in Telegram;
- the assistant downloads the recording and transcribes it;
- the transcript goes into the ordinary pipeline — the same profiles, prompts, FAQ and documents as for text;
- the assistant writes a reply;
- the reply goes out as text or is spoken in the chosen voice — whichever you set.
What matters is that transcription does not run down a separate branch of logic. Once recognised, a voice message is no different from a typed one: the same profiles, the same system prompt and the same knowledge base. There is no separate behaviour to configure for voice.
Transcribing incoming voice notes
Recognition has its own switch. Leave it off and voice messages are ignored exactly as they were before the feature existed.
Transcription is not billed in Stars — it is part of the assistant's work. Only speaking the reply can cost anything, and not always: see below.
Choosing a voice
Voices are shown by ordinary names and a sense of delivery rather than technical codes. The list has male and female voices with different characters — from calm and businesslike to lively and energetic.
They fall into two groups:
- included in your plan — good synthesis, enough for most jobs;
- premium (marked with a star) — the most natural intonation, closest to live speech.
Speed, pitch and volume
Three sliders adjust how the chosen voice sounds:
- speech speed — from noticeably slow to fast;
- pitch — above or below the base;
- volume — quieter or louder.
Premium voices have no pitch control — they are synthesised with an intonation of their own, and the height cannot be changed. In that case the pitch slider is simply not shown, while speed and volume work as usual.
After any change, press Play — the sample is re-synthesised with the new settings and you hear the result straight away.
Reply length and what it does to the price
A separate slider sets the maximum reply length — from 100 to 900 characters. It caps both the text reply and the audio: one lever over both how talkative the assistant is and what the voice costs.
Which gives a simple rule: the shorter the replies, the cheaper the audio. A typical reply from the assistant comes to around 300 characters and costs noticeably less than the 900-character ceiling. The price per reply recalculates as you drag the slider, so you can find the balance between detail and spend by eye, without sending anything.
If a reply does hit the limit, it is trimmed at a sentence boundary rather than mid-word.
What the plan includes
Every plan includes a set number of voice replies a month. Until that runs out, speaking with the included voices is free.
| Plan | Messages a month | Of those, by voice |
|---|---|---|
| Free | 50 | 10 |
| Go | 500 | 100 |
| Plus | 3,000 | 600 |
| Pro | 15,000 | 3,000 |
The remainder is visible right in the settings: how many spoken replies are left from the monthly allowance. The counter resets every month.
Paying beyond the plan with Stars
Anything the plan does not cover is paid for with Telegram Stars from a prepaid balance. That happens in two cases:
- an included voice after the monthly allowance runs out — a very small amount per reply;
- a premium voice — always, and those replies never spend the monthly allowance.
The balance is topped up in bundles right in the mini app, through ordinary Star payments. The order of charges and the current balance are shown there in their own block: how many Stars you have, how many replies that covers in the chosen voice, and what a typical and a maximum reply will cost.
Setting it up in five steps
- Open @Head_Hack_bot and launch the mini app.
- Go to “Profile → Voice replies”.
- Switch the voice side on and leave incoming transcription enabled.
- Pick a reply mode and a voice — listen to a few with the Play button, it is free.
- Set the maximum reply length and save.
After that, send the bot a voice message yourself and check that the transcript is right and the reply sounds the way you expected. This is worth adding to the general pre-launch checklist.
Limitations
- An incoming voice note is recognised if it is under a minute.
- Recognition and synthesis are tuned for Russian.
- The owner's voice is not cloned: the ready neural voices from the list are used instead. That is a deliberate choice — cloning requires separate consent and an identity check.
- Premium voices have no pitch control — only speed and volume.
Common questions
Do I need Telegram Premium?
No. The voice features work like the rest of the assistant — without Premium, without your own bot through BotFather and without a separate server.
Do I pay for transcribing incoming notes?
No. Recognising incoming voice notes is part of the assistant's work. Stars are only charged for speaking replies — and only beyond the monthly allowance or for premium voices.
Can I hear a voice before paying?
Yes, previewing is free for any voice, premium included. You can work through the whole list, change speed and volume and only then decide.
Will the bot speak in my voice?
No. Voice cloning is not used — only the ready voices from the list. You can pick one that suits your style and tune how it sounds.
How do I make the audio cheaper?
Lower the maximum reply length: you are charged for the characters actually spoken, so short replies cost less. The voices included in your plan are several times cheaper than premium.
What happens if someone sends a long voice note?
If the recording is longer than a minute the assistant does not disappear: it asks for something shorter, or for text.
In short
The voice side closes the gap between how people like to write and what the assistant can handle. Incoming voice notes are transcribed free, the reply comes back by default in the format you were addressed in, and the cost of speaking is governed by two understandable levers: the voice you pick and the maximum reply length.
Open @Head_Hack_bot, switch voice replies on, listen to a couple of voices and send the bot a voice message — the whole thing takes a few minutes.