I wanted my Telegram bot to accept a recording and return text. I also wanted to use the same transcription service from another application without extracting its implementation from a bot script. That requirement led to a separate project, vtt, with one application programming interface (API) for uploaded audio.

The service began with ChatGPT’s private transcription endpoint, using the local credential created by Codex login. My Antigravity reverse-engineering experiment supplied the basis for another route: submitting audio to Gemini through Antigravity’s authenticated tooling. I put both behind POST /v1/audio/transcriptions.

My Telegram bot sends an audio file and an engine-specific model identifier. It receives text. Conversion, provider authentication, and response validation belong to the service.

Telegram bot -----+
                  |
Other clients ----+--> My transcription endpoint
                                |
                 +--------------+---------------+
                 |                              |
                 v                              v
       ChatGPT adapter                  Gemini adapter
       split audio as needed            normalize audio once
       private transcription            pinned agy client
                 |                              |
                 +--------------+---------------+
                                |
                                v
                       transcript + engine label

A small contract for the client

I used familiar OpenAI-style routes and response shapes so clients could supply a custom base address. Compatibility here is specific: the service accepts the fields it implements and returns the structures those clients need. It does not reproduce every feature of a hosted API.

Model identifier Behavior
chatgpt-transcribe Transcribe through the ChatGPT adapter.
chatgpt-transcribe-polished Transcribe, then run a separate wording cleanup.
gemini-vtt-transcribe Transcribe through the Antigravity-backed Gemini adapter.

The upload is a multipart request containing file, model, and an optional response_format. The default response is JSON (JavaScript Object Notation) with a text field. A schematic successful response looks like this; the sentence is illustrative:

HTTP/1.1 200 OK
Content-Type: application/json
X-VTT-Engine: gemini

{"text":"I need to revise the introduction."}

HTTP means Hypertext Transfer Protocol. I added X-VTT-Engine so a client could check which adapter answered. A model name in the request proves only what the client asked for.

The service also returns plain text or verbose_json. The latter includes a single compatibility segment spanning the measured recording duration. It provides no word-level timing; those segment fields must not be mistaken for measured speech alignment.

Gemini is explicitly enabled in configuration. Before listing it under /v1/models, the service checks that the required agy executable and model are available. It repeats that check when handling a Gemini transcription. An unavailable Gemini route returns a service-unavailable error. It does not quietly substitute ChatGPT.

There is a narrower compatibility concession: unknown model identifiers use raw ChatGPT transcription, accommodating clients that insist on names such as whisper-1. Consequently, the exact Gemini identifier and the response’s engine header both matter. A misspelled model name is an especially poor engine-selection mechanism.

A successful request can still lose speech

The hardest defect in the ChatGPT path was silent truncation. The project’s recorded probes found responses ending around 9,876 characters despite an HTTP success status. Long audio could produce plausible text with the end missing.

I addressed that in the audio adapter. Its default target is 300 seconds per chunk; it moves boundaries near detected silence and joins the returned text in order. A response of at least 9,000 characters raises suspicion when its source span exceeds 90 seconds. The service splits that span in half and transcribes the halves again, with recovery bounded to three recursive levels.

Recording
   |
   v
Plan spans near silence
   |
   v
Transcribe each span
   |
   +-- suspiciously long output --> halve span --> retry
   |
   v
Join text in recording order

This is a heuristic, a rule that detects a likely failure without proving it. Shorter omissions can escape it, and exhausting the split depth does not certify completeness. The recorded thirteen-minute probe retained 111 of 112 unique markers and produced 13,476 characters. That demonstrated output beyond the observed single-request ceiling; the missing marker also made an unqualified accuracy claim untenable.

Those numbers describe the project’s earlier measurements, not published provider limits or a fresh benchmark. The corresponding code preserves the recovery mechanism because transport success alone says very little about transcription coverage.

Giving Antigravity one bounded job

My first Antigravity client spoke directly to its local service. The reusable endpoint takes a different implementation route: it invokes agy, a command-line client, and validates its structured event stream. It does not execute or import antspoof.py.

The adapter currently requires agy version 1.1.20 and the model label gemini-3.7-flash-low. These are the implementation’s compatibility pins. A different installed version or an absent model makes the adapter unavailable until I deliberately update that contract.

For each request, I create a private temporary directory and convert the recording with FFmpeg, a media-conversion utility. The output is mono WebM with Opus audio, sampled at 48 kilohertz and encoded at 64 kilobits per second. I enforce a 20-mebibyte converted-file cap; one mebibyte is 1,048,576 bytes. This is the adapter’s configured limit, not a general claim about Gemini’s capacity.

The Gemini path submits that converted recording in one run. It does not reuse the ChatGPT chunking algorithm. Larger converted input fails explicitly instead of being divided by an untested second mechanism.

The child process starts in that temporary directory, receives a restricted set of environment variables, and runs with --sandbox and --disable-slash-commands. The prompt tells it to treat the recording as data and return every spoken word in a required transcript field. Spoken instructions are material to transcribe; they confer no authority to perform another task.

I also inspect the returned events. The parser requires the expected workspace, model, conversation, and output schema. It accepts one completed view_file operation targeting the exact audio file. Unexpected tools, subagent activity, malformed events, or a missing successful final result cause rejection. Only the structured transcript becomes the response body.

This inspection occurs after the process runs; rejecting an event stream cannot undo an action. The sandbox and process restrictions serve a separate purpose from validating the returned answer. I do not treat a restrictive prompt as an operating-system boundary.

The default outer deadline is 600 seconds, including conversion. The client’s own deadline is slightly shorter to leave room for cleanup. Output buffers are bounded, and timeout or cancellation triggers process-group cleanup. A transcription request should have a finite lifetime even when its provider has other plans.

The endpoint owns the operational details

The listener binds to 127.0.0.1, the machine’s loopback address. The documented deployment uses a Raspberry Pi and Tailscale Serve for access within my private device network. That describes how I expose the service; it does not mean speech processing happens locally. The selected upstream provider receives the recording.

Clients authenticate with service-specific bearer keys, credentials presented with each request. They do not receive the provider’s login material. Apart from the root health route, requests pass authentication and a physical request-body cap before routing. ChatGPT and Gemini transcription share one concurrency limit, so changing engines cannot bypass admission control.

For the ChatGPT adapter, credential refresh belongs here too. Concurrent authentication failures coordinate through one refresh operation, and replacement credentials are written atomically. Keeping that logic out of Telegram avoids giving every client its own interpretation of token rotation.

The service also exposes text cleanup and speech synthesis. My bot uses cleanup as a separate pass so I can retain the raw transcript. The transcription route remains the central contract: audio and a model identifier in; text and an engine label out.

Private provider interfaces remain version-sensitive dependencies. My own endpoint gives me one place to contain that uncertainty and one contract for clients to use. In the Telegram project, that contract becomes a voice note, a transcript, and a few controls for deciding what to do next.