Agent-readable docs index: /llms.txt. Full docs in one file: /llms-full.txt. Download /docs.zip to grep all markdown files locally.

Batch Mode

Run hundreds of questions against the same context. Cache is always enabled — the first question pays full cost, subsequent questions benefit from provider-level caching at up to 90% savings.

End-to-end example

Audit a TypeScript source tree for security smells in three commands:
End-to-end
# 1. Write the questions cat > audit.txt << 'EOF' Are any credentials hardcoded? Where is user input validated? Are SQL queries parameterized? EOF # 2. Warm the cache (optional: otherwise the first batch question pays full cost) mikro cache --context ./src/ --ext .ts,.js # 3. Run the batch against the warm cache mikro batch audit.txt --context ./src/ --ext .ts,.js --max-cost 1.00
stdout
{"question":"Are any credentials hardcoded?","answer":"No hardcoded credentials detected in src/. All secrets are pulled from process.env via src/config/env.ts.","stats":{"iterations":2,"inputTokens":1500,"outputTokens":220,"cost":0.0070}} {"question":"Where is user input validated?","answer":"Input validation lives in src/middleware/validate.ts using zod schemas per route...","stats":{"iterations":2,"inputTokens":1400,"outputTokens":310,"cost":0.0008}} {"question":"Are SQL queries parameterized?","answer":"All queries in src/db/ use parameterized statements via the pg client; no string interpolation found.","stats":{"iterations":1,"inputTokens":1200,"outputTokens":180,"cost":0.0006}} {"type":"aggregate","total_questions":3,"completed":3,"total_cost":0.0084,"cache_savings":0.0128}
The first question paid ~$0.0070 (cache miss); the next two paid ~$0.0007 each thanks to Gemini's 90% cache discount.

Quick start

Create a questions file (one question per line):
questions.txt
What authentication methods are supported? How does the rate limiter work? What database migrations exist? # This is a comment — skipped Where are the API routes defined?
Run it:
mikro batch questions.txt --context ./src/
Output is JSONL — one JSON object per question, plus a final aggregate line:
{"question":"What authentication methods are supported?","answer":"JWT and OAuth2...","stats":{"iterations":2,"inputTokens":1500,"outputTokens":800,"cost":0.0045}} {"question":"How does the rate limiter work?","answer":"Token bucket algorithm...","stats":{"iterations":3,"inputTokens":1200,"outputTokens":600,"cost":0.0008}} {"question":"What database migrations exist?","answer":"12 migrations in src/db/...","stats":{"iterations":2,"inputTokens":1100,"outputTokens":500,"cost":0.0007}} {"question":"Where are the API routes defined?","answer":"src/routes/ directory...","stats":{"iterations":1,"inputTokens":900,"outputTokens":400,"cost":0.0005}} {"type":"aggregate","total_questions":4,"completed":4,"total_cost":0.0065,"cache_savings":0.0152}

Questions file format

  • One question per line
  • Empty lines are skipped
  • Lines starting with # are treated as comments
What is the project structure? How does error handling work? # Security section What input validation exists? Are there any SQL injection risks?

Cache behavior

Batch mode always enables caching. Here's the cost flow:
QuestionCache statusCost
FirstCache miss (cold)Full input token cost
Second+Cache hit (warm)50-90% cheaper (only cache-read tokens billed)
The exact savings depend on your provider:
ProviderCache discount
Google Gemini~90% on cached input tokens
Anthropic~90% on cached input tokens
OpenAI~50% on cached input tokens

Pre-warming the cache

Warm the cache before running batch queries to ensure the first question also gets cache pricing:
# Warm the cache mikro cache --context ./docs/ # Now run batch — all questions hit warm cache mikro batch questions.txt --context ./docs/

Estimating costs

Check how much a batch run will cost before committing:
mikro cache --context ./docs/ --estimate
mikro cache estimate --- context: ./docs/ metadata: Context is a list of 23 items with 144700 total characters, chunk lengths: [5120, 11873, 2960, ...] estimated tokens: 43,500 provider limit: 1,000,000 tokens utilization: 4.3% provider: google model: gemini-3.1-flash-lite-preview ttl: 3600s estimated cost: $0.0033
estimated cost is the input cost of sending the context once, uncached. mikro prints no cached figure; with Gemini's ~90% cache discount, each later question costs about a tenth of it.
For a 100-question batch over this context: ~$0.003 (first) + 99 * $0.0003 = ~$0.033 total.

Budget enforcement

Set a maximum spend to prevent runaway costs:
mikro batch questions.txt --context ./src/ --max-cost 1.00
Cumulative cost is tracked across all questions. When the budget is exceeded, mikro stops gracefully and reports how many questions were completed.

Batch options

FlagDefaultDescription
--context <path>—Context directory or file
--max-iterations <n>30Max RLM iterations per question
--max-cost <n>—Max total USD spend
--parallel <n>1Accepted; questions currently run one at a time
--batch-apifalseAccepted; mikro does not submit Gemini Batch API jobs yet
--verbosefalseShow progress
Batch output is always JSONL on stdout, so --output has no effect here. The CLI reference lists every batch flag.

Gemini Batch API

mikro batch accepts --batch-api, but mikro does not submit Gemini Batch API jobs yet. A run with the flag makes the same standard calls as a run without it, and no batch discount applies.

Input cost with caching

ModeInput cost (per 1M tokens)Savings
Base (flash-lite)$0.075—
+ Context caching~$0.007590%

Practical patterns

Study session over documentation

# Prepare questions cat > study.txt << 'EOF' What are the core abstractions? How does the plugin system work? What are the extension points? How is state managed? What patterns does the codebase use? EOF # Warm cache, then batch mikro cache --context ./docs/ mikro batch study.txt --context ./docs/ --max-iterations 5

Code audit

cat > audit.txt << 'EOF' Are there any hardcoded credentials? What input validation exists? How are SQL queries constructed? Are there any command injection risks? How is authentication implemented? What error information is leaked to users? EOF mikro batch audit.txt --context ./src/ --ext .ts,.js --max-cost 2.00

Codebase onboarding

cat > onboard.txt << 'EOF' What is the project structure and architecture? What are the main entry points? How is the database accessed? What external services are called? How are tests organized? What CI/CD pipeline is used? EOF mikro batch onboard.txt --context . --ext .ts,.js,.json,.yaml --tools standard

CAG vs RLM for batch

ApproachWhen to use
Batch + cache (CAG)Context fits in provider window, many questions, cost matters
Batch + RLM (no cache)Context too large for system prompt, complex navigation needed
By default, batch mode uses CAG (cache enabled). For very large contexts that exceed provider limits, mikro falls back to standard RLM iteration automatically.

Provider context limits

ProviderMax context (cached)
Google Gemini1,000,000 tokens
Anthropic200,000 tokens
OpenAI128,000 tokens
Amazon Bedrock128,000 tokens
If your context exceeds the provider limit, mikro will warn you and fall back to RLM mode where the LLM navigates the context programmatically.