I Send Voice Notes; My Telegram Bot Sends Back Text
Two transcription engines, durable retries, and a route from my phone to my clipboard.
I wanted to dictate into Telegram and get text I could use elsewhere. The recording might become a paragraph, a note, or material for another project. I wanted the transcript back in the same chat, with an option to send it to a computer’s clipboard. I had little interest in adding another recording interface to my life.
I built the voice bot in my telegram project and kept transcription in a separate vtt service. The bot handles messages, settings, retained recordings, and delivery. My voice-to-text endpoint handles audio conversion and the selected provider.
That separation also explains why I reverse-engineered Antigravity. I wanted its Gemini audio capability available through my own endpoint, where Telegram could use it. The bot never needs to know Antigravity’s internal message fields or local authentication protocol.
My phone
|
| voice note
v
Telegram --> My bot --> My transcription endpoint
^ |
| +--> ChatGPT adapter
| |
| +--> Gemini via Antigravity
| |
+---- transcript --+
|
+--> Telegram text or .txt file
|
+--> Selected computer's clipboard
The clipboard step is optional and follows Telegram delivery. A computer being asleep should not prevent a transcript from reaching my phone.
One recording, several independent choices
I separated four settings: engine, wording, delivery, and clipboard destination. They answer different questions. Which service should hear the recording? Should I retain the spoken wording or clean it up? Should the result arrive as messages or a file? Where else should I copy it?
| Choice | Typed commands | Meaning |
|---|---|---|
| Engine | /codex, /gemini |
Select the ChatGPT or Antigravity-backed Gemini path. |
| Wording | /raw, /polish |
Retain raw wording or request cleanup. |
| Delivery | /auto, /text, /file |
Choose automatic delivery, text messages, or a text file. |
| Clipboard | /copy, /nocopy |
Enable the default destination or disable automatic copying. |
I expose settings through /settings buttons as well as typed commands. Menus make the choices discoverable; commands let me make a known change directly. A caption such as gemini polish file overrides the corresponding settings for one recording.
The defaults are Codex, raw wording, automatic delivery, and no clipboard copying. Per-chat settings persist in SQLite, an embedded relational database. The bot reads settings at the stage where it needs them. Once transcription has started, changing delivery can still send that result as a file; changing the engine does not retroactively change the provider already processing the audio.
Raw and polished wording are separate results. Cleanup asks for punctuation, likely recognition corrections, and removal of fillers without adding content or answering the recording. I retain the raw text, because a cleaner sentence can also be a less faithful one. Polishing is performed through the service’s separate text-cleanup route, including when Gemini produced the original transcript.
The selected engine has to identify itself
The bot calls an application programming interface (API) with one of two model identifiers: chatgpt-transcribe or gemini-vtt-transcribe. Before a Gemini attempt, it checks the service’s model list. If Gemini is unavailable when I select it, the bot reports the failure and leaves the saved engine unchanged.
I also require a matching X-VTT-Engine: gemini response header. A request containing the word Gemini is insufficient evidence that Gemini did the work. This check matters because the endpoint retains an older compatibility behavior for unknown model names.
If the selected engine fails, the bot reports that failure. Switching providers is an explicit retry choice. I want to know which provider produced a transcript, particularly when I am comparing their treatment of the same speech.
The details of the Gemini adapter, including its exact client and model checks, belong in the endpoint article. Telegram’s responsibility is to request the engine and verify the label on the result.
A transcript needs to survive a restart
Once I added actions to delivered transcripts, a temporary in-memory object was insufficient. A button can remain in a Telegram chat long after the process that created it has restarted. I needed durable records behind those controls.
The store separates four kinds of state:
chat_settings: defaults for future decisions
sources: retained recording
|
+--> attempts: engine, status, raw text, polished text
|
+--> views: delivered message and its chosen settings
A source is the recording. An attempt is one provider run. A view records how a particular attempt appeared in Telegram. The transcript lives with the attempt; changing its presentation does not require another transcription or another copy of the canonical text.
That model makes retries useful. I can choose an engine, wording, delivery, and clipboard destination for a new result. Unless I explicitly force another provider call, the bot can reuse the newest successful attempt for that recording and engine. A ChatGPT result cannot satisfy a Gemini retry merely because both contain text.
I retain source recordings for a seven-day window measured from their Telegram upload time. Archive filenames are derived from the chat and message identifiers rather than the uploaded filename. Disk quotas and a free-space reserve constrain storage. Retention gives retries something to work from; it also means this bot keeps audio on disk. It is not an ephemeral transcription relay.
Expiration removes the local source material and ends its retry window. It does not promise deletion of messages already delivered to Telegram or of a provider’s own records.
Buttons come after saved state
The order of operations matters more than the appearance of the keyboard. I save a pending attempt before contacting the provider, then save its success or classified failure before sending the result to Telegram.
Delivery follows another sequence:
Save provider outcome
|
v
Send Telegram message without buttons
|
v
Save message identifier as a view
|
v
Install that message's action buttons
This avoids creating controls whose underlying state was never saved. A crash during recovery can still cause a duplicate message; Telegram delivery and the local database are separate systems, and I cannot commit them in one transaction. I can, however, choose which inconsistency my recovery procedure permits.
The retry buttons encode their choices compactly, and handlers bind the originating message to its stored view. I validate the callback data and recheck chat authorization. A button is another request entry point, even when it looks like an innocent extension of an earlier message.
Getting the text somewhere useful
Automatic delivery sends a short transcript as text and a long one as a .txt attachment. If I request text explicitly, the shared sender splits it at clean boundaries. Telegram documents a 4,096-character limit for an ordinary text message in its Bot API reference.
Those output divisions solve a different problem from the endpoint’s audio chunks. The endpoint divides some recordings to reduce provider truncation. The bot divides returned text to fit Telegram messages. One operation cannot substitute for the other.
Delivery also has failure paths. A failed document upload falls back to inline text. Incomplete inline delivery triggers an attempt to send the full transcript as a document. An empty transcript becomes a visible no-speech placeholder. I would rather receive an explicit result than infer meaning from an absent message.
For clipboard delivery, I pass the transcript through a configured command’s standard input. I do not interpolate spoken text into a shell command. Local destinations use an available clipboard utility. Remote destinations use Secure Shell (SSH), with the destination account’s forced command selecting its clipboard tool. The operation has a fifteen-second timeout and reports failure separately from the transcript.
A small part of a larger Telegram project
The repository also contains a machine-control bot and a one-way notification sender. Shared Telegram message models, request handling, polling, command registration, and text splitting live in tlgbot.py. Each bot owns its own policy and workflow.
For voice transcription, that leaves a clear division of responsibility. The endpoint knows the providers. The bot knows which recording, attempt, and delivered message I am acting on. I can inspect the Antigravity protocol work when I need to repair that adapter; ordinary use starts with a voice note and ends with text I can copy, retain, or revise.