Develop against the Vertex AI generation API, offline.
A tested subset of generateContent and streamGenerateContent through the official genai SDK, and Google's custom prediction contract on CloudBurrow's own runtime. What it covers, what it refuses, and which model it runs.
Generation
A subset of generateContent and streamGenerateContent (server-sent events), verified
with google.golang.org/genai v1.71.0 on both of the
SDK's backend paths: Vertex AI v1beta1 and v1beta.
Refused, not ignored
- generation options
- tools
- safety settings
- multimodal input
- multi-turn
- countTokens
Each returns an error naming the field.
No usageMetadata. No silent model substitution. Off by default; enable it with
--local-ai-model.
The exact surface, and why each field is refused, in docs/generation.md
Model provenance
The only runnable model is litert-community/gemma-4-E2B-it
(2.59 GB), a community conversion, labelled as one
everywhere it appears. Every Google-published artifact is gated.
The console playground says so on screen:
This model is a COMMUNITY conversion, not published by Google. Its output is not a Gemini result and is not labelled as one.
Output quality is not claimed.
Runtime
LiteRT-LM, built from Google's source. CPU only; no GPU or Metal.
In v0.1.0 the runtime is built with make litert-lm from a checkout; a prebuilt runtime image comes
with the next release (#602). next release
Custom prediction
Google's serving contract (AIP_HTTP_PORT, the health and predict routes, instances in
and predictions out) on CloudBurrow's own runtime.
Not supported
- Model Registry
- Endpoints
- PredictionService
- Model Garden
- tuning
- batch prediction
AIP_STORAGE_URI
A container that fails to start is reported after 240 seconds. next release
Embeddings
Blocked by the model export today (#41). What was tried, in docs/embeddings.md
Build against Google Cloud APIs, locally
Free and open source under Apache-2.0. No account, no sign-up, no Google Cloud bill.