Google Antigravity SDK Runs Gemma 4 Models Fully Offline on Local Hardware
Google has just made a significant move in the world of on-device AI. The Antigravity SDK now supports running agentic workflows completely offline using Gemma 4 26B models through Google AI Edge's LiteRT runtime. This update means developers can execute sophisticated AI agents directly on their local machines—no cloud API calls, no token costs, and no code ever leaving their hardware.
The announcement, made on the Google for Developers Blog, marks a shift toward local orchestration for privacy-sensitive and cost-conscious development environments.
📋 Table of Contents
What Is the Antigravity SDK Local Model Update?
The Antigravity SDK is Google's framework for building agentic applications—AI systems that can perform multi-step tasks like code auditing, file editing, and autonomous workflows. Previously, these agents required cloud-based models like Gemini.
Now, with this update, developers can configure the SDK to run entirely on-device using the Gemma 4 26B A4B model checkpoint. The execution is powered by LiteRT, Google's on-device runtime (formerly TensorFlow Lite), which handles GPU acceleration automatically on supported hardware.
| Feature | Details |
|---|---|
| Default Local Model | Gemma 4 26B A4B |
| Runtime | Google AI Edge's LiteRT |
| Model Size | Approximately 16.8 GB download |
| Hardware Requirement | >24 GB VRAM or unified memory recommended |
| Alternative Backends | Ollama, LM Studio, vLLM (via OpenAI-compatible API) |
| Key Benefit | Zero token costs, full offline capability |
Why Run AI Agents Locally?
Google outlined several compelling reasons for developers to adopt local model execution:
Cost Efficiency
Agentic workflows are token-hungry. A single code refactoring task can trigger hundreds of model calls, making cloud API costs prohibitive for frequent use. Local execution reduces marginal costs to essentially the price of electricity.
Privacy and Compliance
For teams in finance, healthcare, defense, or any environment with strict data handling requirements, keeping source code entirely on local machines is essential. Local execution ensures no proprietary code ever leaves the developer's hardware.
Offline Resiliency
Once the model is downloaded and configured, the agent works without any internet connection. This is critical for developers in environments with unstable connectivity or strict network restrictions.
Hybrid Workflows
Google demonstrated an "Architect-Builder" pattern where a cloud model (Gemini 3.8 Flash) acts as the planner, decomposing tasks based on filenames and descriptions, while local Gemma 4 instances handle the actual code work.
Hardware Requirements: The 24 GB Threshold
Google recommends a machine with more than 24 GB of VRAM or unified memory to run the Gemma 4 26B A4B checkpoint comfortably.
The 24 GB recommendation isn't arbitrary. The 26B parameter model, even when quantized, consumes significant memory for weights alone. The remaining memory must accommodate the KV cache, context window, and the LiteRT runtime itself.
What This Means for Different Hardware
- Apple Silicon users: MacBook Pro or Mac Studio with 48GB or 64GB unified memory are ideal, as memory is shared between CPU and GPU.
- Windows/Linux with discrete GPU: You'll need something in the RTX 4090 or RTX 5090 tier to meet the 24 GB VRAM threshold.
- Machines with 16 GB RAM: May load the model but will experience severe performance degradation as context grows, with latency dropping from hundreds of milliseconds to several seconds per turn.
Run open models like Gemma 4 completely offline in the Antigravity SDK.
— Google Antigravity (@antigravity) September 23, 2026
Built on Google AI Edge’s LiteRT, you can now run Gemma 4 directly on your local GPU. Zero API costs, total data privacy, and no internet required. pic.twitter.com/15mmQiXF4m
How to Get Started: Step-by-Step Setup
Here is the complete setup process based on Google's official documentation:
Step 1: Create a Virtual Environment
python3 -m venv .venv
source .venv/bin/activate
Step 2: Install the Required Packages
pip install google-antigravity litert-lm
Step 3: Download the Gemma 4 26B Model
litert-lm import \
--from-huggingface-repo=litert-community/gemma-4-26B-A4B-it-litert-lm \
gemma-4-26B-A4B-it-gpu.litertlm \
gemma4-26b
This command downloads approximately 16.8 GB and registers the checkpoint at ~/.litert-lm/models/gemma4-26b/model.litertlm.
certifi and set the environment variable:
pip install certifi
export SSL_CERT_FILE=$(python3 -c "import certifi; print(certifi.where())")
Step 4: Create a Basic Agent Script
Create a file named agy_sample.py:
import asyncio
import os
from google.antigravity import Agent, LiteRTAgentConfig
MODEL_PATH = os.path.expanduser(
"~/.litert-lm/models/gemma4-26b/model.litertlm"
)
async def main():
print(f"Using local LiteRT model: {MODEL_PATH}")
print("Please wait for local inference to complete. This could take several minutes.")
config = LiteRTAgentConfig(model_path=MODEL_PATH)
async with Agent(config) as agent:
response = await agent.chat("What files are in the current directory?")
async for token in response:
print(token, end="", flush=True)
print()
if __name__ == "__main__":
asyncio.run(main())
Step 5: Run the Agent
python agy_sample.py
The LiteRTAgentConfig automatically applies a lightweight preset that restricts built-in tools to essential coding operations, reducing prompt overhead for local hardware.
Alternative: Connecting to Local Server Backends
If you already have a local inference server running, the SDK supports connecting via an OpenAI-compatible API. This works with Ollama, LM Studio, or vLLM:
from google.antigravity import Agent, LocalOpenAIAgentConfig
config = LocalOpenAIAgentConfig(
base_url="http://localhost:11434/v1",
model="gemma-4-26B-A4B-it"
)
This approach lets you switch between different local model backends without changing your agent orchestration, tools, or workflows.
Hybrid Orchestration: Cloud Planner + Local Workers
One of the most compelling demonstrations from Google shows a hybrid workflow where cloud and local models collaborate:
The Setup
- Cloud Architect: Gemini 3.8 Flash receives only filenames and task descriptions (no source code).
- Local Builders: Multiple Gemma 4 26B instances run on the local GPU.
The Task
Audit and patch three vulnerable Python modules (auth.py, billing.py, database.py).
The Workflow
- Gemini plans the strategy and decomposes work (95 cloud tokens spent).
- Local Gemma instances reproduce vulnerabilities, author fixes, critique patches, and validate against regression tests.
- Total token usage: 3,322 tokens, with 97.2% processed locally and offline.
Frequently Asked Questions (FAQs)
Q1: What hardware do I need to run Gemma 4 26B locally?
Google recommends a machine with more than 24 GB of VRAM or unified memory. Apple Silicon Macs with 48GB+ unified memory work well. For Windows/Linux, you'll need a high-end GPU like RTX 4090 or RTX 5090.
Q2: Is this truly free? Are there any hidden costs?
The local execution itself has zero token costs—you're running on your own hardware. The only costs are the one-time hardware investment and electricity. However, if you use the hybrid pattern, any cloud tokens consumed by the planner model would be billed normally.
Q3: Can I use a different model besides Gemma 4 26B?
Yes. The LocalOpenAIAgentConfig lets you connect to any OpenAI-compatible local server. You can run Ollama, LM Studio, or vLLM with other open-weights models. However, the optimized LiteRT path is currently tuned for Gemma 4 26B.
Q4: How long does local inference take?
Google's sample warns that local inference "could take several minutes" for complex tasks. This depends heavily on your hardware. Apple Silicon Macs with unified memory tend to perform well for local LLM inference.
Q5: What's the difference between this and just running Ollama?
The Antigravity SDK provides the agentic orchestration layer—tool calling, file editing, multi-step workflows, and policy controls. Running Ollama alone gives you a model endpoint but not the full agent framework. The SDK can connect to Ollama as a backend while providing all the agent capabilities.
Q6: Can I run this on a mobile device?
The current recommendation is for desktop-class hardware with 24GB+ memory. However, Google has separately demonstrated Gemma 4 12B running on laptops with 16GB RAM, and Gemma 4 E4B on Raspberry Pi devices for specialized use cases.
Q7: What happens if my machine has less than 24 GB of memory?
The model may still load on 16 GB systems, but performance will degrade significantly as context grows. The system will start paging to disk, causing latency to spike from milliseconds to seconds per turn, making agentic workflows impractical.
Q8: Is this ready for production use?
Google's examples demonstrate working prototypes, but the hybrid workflow's reliability across arbitrary codebases hasn't been established. The recorded audit run produced verified patches, but teams should validate performance on their own projects before production deployment.
Key Takeaways
| Feature | Details |
|---|---|
| What's new | Antigravity SDK supports fully offline agentic workflows |
| Default model | Gemma 4 26B A4B via LiteRT |
| Hardware requirement | >24 GB VRAM or unified memory |
| Model size | ~16.8 GB download |
| Cost | Zero token costs for local execution |
| Privacy | Source code never leaves the machine in fully local mode |
| Hybrid mode | 97.2% of tokens processed locally in Google's demo |
| Alternative backends | Ollama, LM Studio, vLLM via OpenAI-compatible API |
Official Links
- Google for Developers Blog Announcement: Introducing Support for Local AI Models in the Antigravity SDK
- Antigravity SDK Documentation: Local Models Guide
- GitHub Example: local_models.py on GitHub
Read Google's Official Announcement → View SDK Documentation →
0 टिप्पणियाँ