| Mode | How it works | Best for |
| Default RLM | Context loaded into a Python REPL context variable; LLM writes code to navigate it programmatically | Very large codebases, exploratory analysis, unknown questions |
| Cache (CAG) | Full context embedded in the system prompt and cached at the provider | Repeated questions on the same docs, study sessions, batch Q&A |
--cache flag on any query, or set cache.enabled: true in mikro.yaml inside .mikro/.1mikro cache --context ./docs/ --estimate
1234567891011mikro cache estimate --- context: ./docs/ metadata: Context is a list of 42 items with 309580 total characters, chunk lengths: [11904, 3127, 8466, ...] estimated tokens: 93,000 provider limit: 1,000,000 tokens utilization: 9.3% provider: google model: gemini-3.1-flash-lite-preview ttl: 3600s estimated cost: $0.0070
1mikro cache --context ./docs/
1234567mikro: warming cache for ./docs/ (~93,000 tokens) mikro: cache warmup complete provider: google model: gemini-3.1-flash-lite-preview estimated tokens: 93,000 ttl: 3600s estimated cost: $0.0070
ttl line is mikro's own display value.123mikro "What RPC primitives are available?" --context ./docs/ --cache mikro "How are errors surfaced?" --context ./docs/ --cache mikro "What's the threading model?" --context ./docs/ --cache
12345Warmup (mikro cache): $0.0070 full input tokens billed Query 1 (cached): $0.0007 90% discount on cached input Query 2 (cached): $0.0007 90% discount on cached input Query 3 (cached): $0.0007 90% discount on cached input Total for four runs: ~$0.0091
--cache run sends the full context in the system prompt again, with a session ID built from the content hash. While the provider still holds that prefix, the repeated input tokens are billed at the cached rate.--cache is enabled mikro:{cache.session-prefix}-{hash} (or just the hash)## Context Files blockcache_control for Anthropic, a prompt cache key for OpenAI); Gemini caches repeated prefixes on its owncache_read_tokens — billed at the discount ratestorage.enabled is auto/always).| Provider | Cache limit | Discount on cached input | TTL behavior |
| Google Gemini | 1,000,000 tokens | ~90% | Implicit caching by Google; mikro sends no cache TTL, so retention has no effect |
| Anthropic | 200,000 tokens | ~90% | Ephemeral (~5 min) or long-lived via cache_control |
| OpenAI | 128,000 tokens | ~50% | Automatic prompt caching; retention: long, mikro's default, asks for 24-hour retention where the model supports it |
| Amazon Bedrock | 128,000 tokens | Provider-dependent | Inherits underlying model support |
cache.retention: long, which asks for the 1-hour tier where the model supports it (usually 2× base cost to write, ~90% discount on reads). Set retention: short in your mikro.yaml for the 5-minute tier.cache: in mikro.yaml inside .mikro/:123456cache: enabled: false # enable globally (or use --cache per-invocation) retention: long # short | long, maps to the provider's cache retention ttl: 3600 # seconds; shown by `mikro cache`, not sent to the provider expire-time: "" # ISO 8601; accepted, not sent to the provider session-prefix: "myproj" # prepended to the content hash in the session ID
| Field | Description |
enabled | Turn cache mode on by default for every mikro invocation. CLI --cache overrides this per-run. |
retention | short for ephemeral caches, long for extended TTL. Maps to provider-specific behavior. |
ttl | Seconds. Shown in mikro cache output; mikro does not send it to the provider. |
expire-time | ISO 8601 timestamp. Accepted, but mikro does not send it to the provider. |
session-prefix | Namespace for the cache session ID — useful when multiple projects share a provider account. |
mikro cache commandmikro cache is the operator-facing entry point for CAG. It has two modes:| Invocation | Behavior |
mikro cache --context <path> --estimate | Prints token count, provider limit, utilization %, and projected first-query cost. No LLM calls. |
mikro cache --context <path> | Issues a one-iteration warmup query to prime the provider cache. |
mikro cache.1234mikro cache --context ./src/ --ext .ts,.js mikro "Where is the auth middleware?" --context ./src/ --cache --ext .ts,.js mikro "What drives rate limiting?" --context ./src/ --cache --ext .ts,.js mikro "List the database migrations" --context ./src/ --cache --ext .ts,.js
mikro cache plus a question loop. See Batch Mode for the questions-file format and cost math.12mikro cache --context ./docs/ mikro batch study.txt --context ./docs/
12345mikro "Summarize the entire repo" \ --context ./src/ \ --cache \ --max-cost 0.50 \ --max-iterations 10
12mikro: context exceeds model limit (~1,250,000 tokens > 1,000,000), disabling cache mode mikro: storage mode activated for large context (~1,250,000 tokens)
storage.enabled: auto)