Agent-readable docs index: /llms.txt. Full docs in one file: /llms-full.txt. Download /docs.zip to grep all markdown files locally.

Cache Mode (CAG)

Cache-Augmented Generation (CAG) embeds your full context directly in the system prompt and lets the provider's prompt cache handle the rest. Subsequent queries over the same context pay a fraction of the input-token cost — up to 90% off on Google and Anthropic.
Use cache mode when your context fits inside the provider's window and you plan to ask multiple questions against it. For corpora that exceed the window, stay on default RLM mode (programmatic navigation).

CAG vs default RLM

ModeHow it worksBest for
Default RLMContext loaded into a Python REPL context variable; LLM writes code to navigate it programmaticallyVery large codebases, exploratory analysis, unknown questions
Cache (CAG)Full context embedded in the system prompt and cached at the providerRepeated questions on the same docs, study sessions, batch Q&A
Switch into cache mode with the --cache flag on any query, or set cache.enabled: true in mikro.yaml inside .mikro/.

End-to-end example

Interrogate a documentation folder three times: one warmup call fills the cache, and the three queries after it ride it.

1. Estimate the cost

mikro cache --context ./docs/ --estimate
Output
mikro cache estimate --- context: ./docs/ metadata: Context is a list of 42 items with 309580 total characters, chunk lengths: [11904, 3127, 8466, ...] estimated tokens: 93,000 provider limit: 1,000,000 tokens utilization: 9.3% provider: google model: gemini-3.1-flash-lite-preview ttl: 3600s estimated cost: $0.0070
The context fits well under the 1M-token Gemini window — safe to cache.

2. Warm the cache

mikro cache --context ./docs/
Output (stderr)
mikro: warming cache for ./docs/ (~93,000 tokens) mikro: cache warmup complete provider: google model: gemini-3.1-flash-lite-preview estimated tokens: 93,000 ttl: 3600s estimated cost: $0.0070
One warmup call primes Gemini's implicit cache. Google decides how long it lasts; the ttl line is mikro's own display value.

3. Run cached queries

mikro "What RPC primitives are available?" --context ./docs/ --cache mikro "How are errors surfaced?" --context ./docs/ --cache mikro "What's the threading model?" --context ./docs/ --cache
Approximate input cost per run
Warmup (mikro cache): $0.0070 full input tokens billed Query 1 (cached): $0.0007 90% discount on cached input Query 2 (cached): $0.0007 90% discount on cached input Query 3 (cached): $0.0007 90% discount on cached input Total for four runs: ~$0.0091
Each --cache run sends the full context in the system prompt again, with a session ID built from the content hash. While the provider still holds that prefix, the repeated input tokens are billed at the cached rate.

How it works

When --cache is enabled mikro:
  1. Computes a SHA-256 content hash over sorted context items
  2. Builds a session ID: {cache.session-prefix}-{hash} (or just the hash)
  3. Writes the full context into the system prompt under a ## Context Files block
  4. Sends the request with the provider's cache options (cache_control for Anthropic, a prompt cache key for OpenAI); Gemini caches repeated prefixes on its own
  5. Provider returns usage metrics including cache_read_tokens — billed at the discount rate
If the context exceeds the provider limit, mikro logs a warning, disables cache mode automatically, and falls back to RLM navigation (or storage mode if storage.enabled is auto/always).

Provider support

ProviderCache limitDiscount on cached inputTTL behavior
Google Gemini1,000,000 tokens~90%Implicit caching by Google; mikro sends no cache TTL, so retention has no effect
Anthropic200,000 tokens~90%Ephemeral (~5 min) or long-lived via cache_control
OpenAI128,000 tokens~50%Automatic prompt caching; retention: long, mikro's default, asks for 24-hour retention where the model supports it
Amazon Bedrock128,000 tokensProvider-dependentInherits underlying model support
Anthropic note: Anthropic's default cache TTL is ~5 minutes. mikro defaults to cache.retention: long, which asks for the 1-hour tier where the model supports it (usually 2× base cost to write, ~90% discount on reads). Set retention: short in your mikro.yaml for the 5-minute tier.

Configuration

Cache behavior lives under cache: in mikro.yaml inside .mikro/:
cache: enabled: false # enable globally (or use --cache per-invocation) retention: long # short | long, maps to the provider's cache retention ttl: 3600 # seconds; shown by `mikro cache`, not sent to the provider expire-time: "" # ISO 8601; accepted, not sent to the provider session-prefix: "myproj" # prepended to the content hash in the session ID
FieldDescription
enabledTurn cache mode on by default for every mikro invocation. CLI --cache overrides this per-run.
retentionshort for ephemeral caches, long for extended TTL. Maps to provider-specific behavior.
ttlSeconds. Shown in mikro cache output; mikro does not send it to the provider.
expire-timeISO 8601 timestamp. Accepted, but mikro does not send it to the provider.
session-prefixNamespace for the cache session ID — useful when multiple projects share a provider account.
See the full table in Configuration → cache.

The mikro cache command

mikro cache is the operator-facing entry point for CAG. It has two modes:
InvocationBehavior
mikro cache --context <path> --estimatePrints token count, provider limit, utilization %, and projected first-query cost. No LLM calls.
mikro cache --context <path>Issues a one-iteration warmup query to prime the provider cache.
Full flag reference: CLI → mikro cache.

Practical patterns

Study session over a codebase

Warm once, then ask as many follow-ups as you want within the TTL:
mikro cache --context ./src/ --ext .ts,.js mikro "Where is the auth middleware?" --context ./src/ --cache --ext .ts,.js mikro "What drives rate limiting?" --context ./src/ --cache --ext .ts,.js mikro "List the database migrations" --context ./src/ --cache --ext .ts,.js

Cached batch interrogation

Batch mode always enables caching, so it works like mikro cache plus a question loop. See Batch Mode for the questions-file format and cost math.
mikro cache --context ./docs/ mikro batch study.txt --context ./docs/

Budget-capped cached run

Set a hard spend ceiling so runaway iterations can't overshoot cache savings:
mikro "Summarize the entire repo" \ --context ./src/ \ --cache \ --max-cost 0.50 \ --max-iterations 10

Automatic fallback to RLM

If the context inflates past the provider limit, mikro logs a warning and downgrades to RLM navigation:
mikro: context exceeds model limit (~1,250,000 tokens > 1,000,000), disabling cache mode mikro: storage mode activated for large context (~1,250,000 tokens)
No action needed — the query still runs, just without caching.

When NOT to use cache mode

  • Context is too large for the provider window (use default RLM with storage.enabled: auto)
  • Questions span different contexts — each unique context pays its own warmup cost
  • You only plan to ask one question — the first-query cost equals a non-cached query, so there's no savings

See also

Batch Mode
Bulk interrogation over one cached context, with cost estimation up front.
CLI Reference
Every flag on mikro cache and --cache documented.
Configuration
cache: section of mikro.yaml in full.
Provider limits
Max cacheable context size by provider.