I Reverse-Engineered Antigravity for My Voice Bot
A desktop application’s internal protocol became a way to transcribe my Telegram recordings.
I wanted to send a voice note to my Telegram bot and get usable text back. I already had a separate voice-to-text service; I wanted to use Antigravity’s audio capability through my own application programming interface (API), with Telegram as the client. Opening a code editor every time I wanted a transcript was a peculiar requirement for dictation.
That practical problem sent me into Google’s Antigravity integrated development environment (IDE). I needed to understand what its microphone submitted, which process handled the request, and whether my own program could reproduce the exchange.
I expected an intermediary that converted speech to text before the model request. The client code showed a different sequence. The chat composer attached the recording itself. It packaged Opus-encoded audio in a WebM container, encoded the bytes as base64 text, and submitted them as message media. I found no separate speech-to-text call in that client path.
Following that request led me through minified JavaScript, a 131-megabyte Go binary, and an embedded service schema. I ended up with a small Python client that could send a twenty-minute recording through Antigravity’s signed-in local service. I supplied no separate Google API key and used no software development kit (SDK). The application still had to be running; its authentication was doing substantial work.
This article covers the reverse engineering. I describe the resulting service design in One Endpoint for My Voice Notes, and the interface I actually wanted in I Send Voice Notes; My Telegram Bot Sends Back Text.
These are observations from the installation I inspected. Internal method names, model aliases, and authentication behavior belong to that build; they are not a compatibility promise from Google.
The local service owns the request
Antigravity uses Electron and Visual Studio Code components. In the chat path I traced, the interface sent its model requests to a local background process, or daemon, named language_server. That process handled the upstream connection and the signed-in session.
The request crossed three components in sequence: the chat interface, the daemon listening on the loopback address 127.0.0.1, and Google’s service. Loopback confines the network connection to the same machine. It does not, by itself, distinguish one local program from another.
That architecture gave me a narrower problem. I needed to reproduce the interface’s requests to the daemon. I did not need to implement Google’s upstream authentication flow.
Antigravity chat -----+
|
v
Local language_server ---> Google
^ |
| |
My Python client ----+ <-----------+
audio + prompt response
Both clients used the local service. My experiment replaced the request’s origin while retaining the signed-in daemon.
Inside the application bundle, I found the daemon at extensions/antigravity/bin/language_server_macos_arm. The minified panel and extension JavaScript supplied the transport details. Ordinary calls sent JSON (JavaScript Object Notation) with an x-codeium-csrf-token header; streaming calls used application/connect+json.
The launch arguments supplied more context:
language_server --standalone
--override_ide_name antigravity
--https_server_port 0
--csrf_token <session-token>
--api_server_url https://generativelanguage.googleapis.com
--cloud_code_endpoint https://daily-cloudcode-pa.googleapis.com
--enable_sidecars
I have replaced the token with a placeholder. The significant detail is its location: a process argument. In my session, the process listing exposed it, and the open-file listing identified the listening ports. A call to GetStatus, carrying that token, returned an empty successful response.
The token’s name refers to cross-site request forgery (CSRF), in which a website induces a browser to send an unwanted request. A defense against that threat does not necessarily isolate a service from programs running under the same user account. My client operated within that local account boundary.
The binary supplied the schema
The most useful artifact was already inside the executable. Protocol Buffers, Google’s schema-based serialization system, supports descriptors that describe message fields and services. The Go implementation can retain those descriptors for runtime reflection, which lets a program inspect its own types. Google’s Go generated-code documentation describes this reflection machinery.
I searched the binary for exa.language_server_pb and located embedded descriptor data. A short Python parser recovered 1,359 message definitions and 290 remote procedure calls (RPCs), meaning named operations exposed by the service.
My parser handled the wire types needed for that extraction, including variable-length integers and length-delimited fields. Those are not the entire Protocol Buffers wire format; the encoding specification also defines fixed-width fields and legacy groups. A useful parser need not become an unsolicited serialization framework.
The recovered definitions included these fields; I have omitted unrelated fields:
message Media {
string mime_type = 1;
bytes inline_data = 2;
float duration_seconds = 7;
}
message SendUserCascadeMessageRequest {
string cascade_id = 1;
repeated TextOrScopeItem items = 2;
Metadata metadata = 3;
CascadeConfig cascade_config = 5;
repeated Media media = 14;
}
Here, cascade_id identifies the conversation, items carries text, and media carries attachments. The numbered fields are serialization identifiers. In the JSON representation, inline_data becomes inlineData, and its bytes become base64 text, consistent with the Protocol Buffers JSON mapping.
The runtime catalog supplied the next piece of evidence. GetAvailableModels returned 28 models in my session. Every Gemini chat entry I inspected advertised audio/webm;codecs=opus among its supported media types. MIME (Multipurpose Internet Mail Extensions) types identify a payload’s format; this one declares a WebM container with Opus audio.
Together, the client code, message schema, and model catalog supported the same conclusion: the chat interface submitted audio as an attachment to the model request. They did not reveal every operation Google might perform after receiving it. Claims about the absence of all server-side transcription would exceed the evidence.
The errors filled in the remaining fields
The schema made requests possible. The daemon’s errors made them usable.
An unsupported-media-type response told me that ordinary calls expected application/json. A missing-source error identified a required conversation enum, an enumerated value selected from a fixed set: CORTEX_TRAJECTORY_SOURCE_INTERACTIVE_CASCADE. Another error led to the model configuration under cascadeConfig.plannerConfig.requestedModel.
The authentication error was less informative. A request with the correct token failed until I included metadata.ideName: "antigravity". That established a metadata requirement for the call I tested. It did not establish the interceptor’s complete validation logic.
Model selection required a runtime lookup. The catalog mapped friendly names to identifiers such as MODEL_PLACEHOLDER_M301; the binary contained more than 650 placeholder values. In my run, the catalog associated that example with gemini-3.7-flash-tiered.
I treated those strings as internal catalog labels. The prototype queried the catalog when I explicitly selected a model, but retained hard-coded defaults and a fallback identifier. That was sufficient for the experiment and unsuitable as a durable contract. The later service adapter checks its exact client version and model before advertising Gemini availability.
A small protocol with consequential details
The daemon used Connect, an RPC protocol built on HTTP (Hypertext Transfer Protocol). Its JSON encoding made a generated client optional.
An ordinary request had this shape:
POST /exa.language_server_pb.LanguageServerService/GetStatus HTTP/1.1
Content-Type: application/json
Connect-Protocol-Version: 1
x-codeium-csrf-token: <session-token>
{}
Streaming messages used five-byte headers: one byte of flags, followed by a four-byte payload length in big-endian order, with the most significant byte first. The JSON payload followed the header. A frame encoder was correspondingly short:
def envelope(payload, flags=0):
return struct.pack(">BI", flags, len(payload)) + payload
Connect reserves flag 0x02 for the response’s end-of-stream envelope. A successful terminal payload can be {}; an empty payload is not a JSON object. For a server-streaming call, the request contains one framed message. The protocol does not require a corresponding terminal envelope on that request.
The local endpoint used a self-signed certificate. My experimental client disabled Transport Layer Security (TLS) certificate verification for that loopback connection. That was a limitation of the client: encryption without certificate verification does not authenticate the peer.
Sending the recording
I converted the input recording to WebM with Opus audio using FFmpeg, a media-conversion utility. That reproduced the container and codec expected by the chat interface. It did not reproduce the browser’s bytes, encoder settings, or metadata. Matching media types offers no guarantee against fingerprinting.
After finding the daemon and checking GetStatus, the client created a conversation with StartCascade. I supplied a fresh conversation identifier, a workspace location, the interactive-cascade source and trajectory values, and the required client metadata.
The subsequent message carried text and audio together. This excerpt shows the essential shape; the model name and enum are examples from my runtime catalog:
{
"cascadeId": "<conversation-id>",
"items": [{"text": "Transcribe the recording verbatim."}],
"media": [{
"mimeType": "audio/webm;codecs=opus",
"inlineData": "<base64-encoded-audio>"
}],
"cascadeConfig": {
"plannerConfig": {
"modelName": "gemini-3.7-flash-tiered",
"google": {},
"requestedModel": {"model": "MODEL_PLACEHOLDER_M301"}
}
},
"metadata": {"ideName": "antigravity", "os": "darwin"}
}
The daemon saved the attachment beneath the conversation’s .user_uploaded directory in ~/.gemini/antigravity/brain/. The WebM file received an .img extension. The extension did not describe the file’s format.
I then polled GetCascadeTrajectory. The returned steps included user input, checkpoints, and planner responses. Once the run stopped, I extracted the final planner response’s response field.
The client used Python’s standard library for networking, serialization, framing, and process discovery. Its external requirements were the operating system’s ps and lsof utilities, FFmpeg, and an open, signed-in Antigravity session. No additional Python packages were required.
Authentication remained the daemon’s responsibility
My requests carried the local token and client metadata. The daemon supplied the upstream credentials from the application’s signed-in session. I did not need to extract or copy those credentials into the script.
That arrangement explains why the client worked without a separate API key. It does not establish that Google could not distinguish its use from ordinary interaction. Request timing, content, client metadata, and service-side checks could all matter. I did not measure detection, and successful requests are not evidence of permission under a service’s terms.
The operational dependency was straightforward: the application had to remain open and signed in. I had replaced the chat interface for these requests; I still depended on the service it launched and the access it held.
Dictation followed a separate path
The editor’s inline dictation used automatic speech recognition (ASR), a service that returns text from speech. Its client path differed from the chat attachment path.
An AudioWorklet, a browser component for processing audio, downsampled microphone input from 48 to 16 kilohertz. It produced mono pulse-code modulation (PCM) samples: signed 16-bit integers in little-endian byte order, with the least significant byte first.
StreamAudioTranscription opened a session. Repeated SendAudioChunk calls delivered 32-kilobyte chunks with sequence numbers. EndAudioSession requested completion. The response schema exposed TranscriptionUpdate records with text and an isFinal flag for partial and final results.
I implemented that exchange, but the backend rejected the session with PermissionDenied: insufficient authentication scopes. Scopes specify the access granted to a credential. The error identified an authorization problem; it did not tell me whether the cause lay in the login flow, account entitlement, or service configuration. I therefore did not count dictation as a successful end-to-end test.
What the recording demonstrated
I began with a three-second clip: Hello Antigravity, this is a test of the voice pipeline.
The model returned those words.
The larger test used an Ogg recording lasting 1,201 seconds, approximately twenty minutes, with a source-file size of 4.4 megabytes. I asked for a verbatim transcript, preserving the order of speech and omitting no spoken content.
The response appeared to cover the full monologue. It retained filler words, repetitions, and mid-sentence corrections. Those features were consistent with the requested transcription rather than a condensed account. They did not establish zero omissions or a measured error rate; I had not performed a word-by-word accuracy audit.
The experiment demonstrated something narrower and useful: my local client could submit a long audio attachment through the authenticated daemon and retrieve a substantial transcription from the selected model. The recording did not have to originate in the application’s microphone control.
From the experiment to my endpoint
The twenty-minute recording answered the question that mattered for my bot: audio could originate in Telegram and still be submitted through Antigravity’s authenticated service. The editor’s microphone was one source of those bytes; my own program could supply another.
I kept the exploratory antspoof.py client separate from the reusable service. In the current voice-to-text implementation, the Gemini adapter invokes the agy command-line client, converts the recording to WebM with Opus audio, and requires a structured transcript. It checks the client version and model, bounds the process lifetime, and validates the returned event stream. The service does not import the experimental script.
That distinction matters. Recovering field names solved the protocol problem. A reusable endpoint also needed predictable failures, temporary-file cleanup, and an answer that belonged to the requested recording. I describe those choices in One Endpoint for My Voice Notes.
My Telegram bot selects that backend with the model identifier gemini-vtt-transcribe. It submits audio to my endpoint and receives text; the endpoint owns the Antigravity-specific work. I can change the transcription adapter without teaching the bot another internal protocol.
The reverse engineering gave me a usable boundary: Antigravity’s interface and its authenticated request service were separate components. I could build my own client around that separation. The result was a path from a voice note to text in the chat where I had wanted it all along.