123456789101112# 1. Write the questions cat > audit.txt << 'EOF' Are any credentials hardcoded? Where is user input validated? Are SQL queries parameterized? EOF # 2. Warm the cache (optional: otherwise the first batch question pays full cost) mikro cache --context ./src/ --ext .ts,.js # 3. Run the batch against the warm cache mikro batch audit.txt --context ./src/ --ext .ts,.js --max-cost 1.00
1234{"question":"Are any credentials hardcoded?","answer":"No hardcoded credentials detected in src/. All secrets are pulled from process.env via src/config/env.ts.","stats":{"iterations":2,"inputTokens":1500,"outputTokens":220,"cost":0.0070}} {"question":"Where is user input validated?","answer":"Input validation lives in src/middleware/validate.ts using zod schemas per route...","stats":{"iterations":2,"inputTokens":1400,"outputTokens":310,"cost":0.0008}} {"question":"Are SQL queries parameterized?","answer":"All queries in src/db/ use parameterized statements via the pg client; no string interpolation found.","stats":{"iterations":1,"inputTokens":1200,"outputTokens":180,"cost":0.0006}} {"type":"aggregate","total_questions":3,"completed":3,"total_cost":0.0084,"cache_savings":0.0128}
12345What authentication methods are supported? How does the rate limiter work? What database migrations exist? # This is a comment — skipped Where are the API routes defined?
1mikro batch questions.txt --context ./src/
12345{"question":"What authentication methods are supported?","answer":"JWT and OAuth2...","stats":{"iterations":2,"inputTokens":1500,"outputTokens":800,"cost":0.0045}} {"question":"How does the rate limiter work?","answer":"Token bucket algorithm...","stats":{"iterations":3,"inputTokens":1200,"outputTokens":600,"cost":0.0008}} {"question":"What database migrations exist?","answer":"12 migrations in src/db/...","stats":{"iterations":2,"inputTokens":1100,"outputTokens":500,"cost":0.0007}} {"question":"Where are the API routes defined?","answer":"src/routes/ directory...","stats":{"iterations":1,"inputTokens":900,"outputTokens":400,"cost":0.0005}} {"type":"aggregate","total_questions":4,"completed":4,"total_cost":0.0065,"cache_savings":0.0152}
# are treated as comments123456What is the project structure? How does error handling work? # Security section What input validation exists? Are there any SQL injection risks?
| Question | Cache status | Cost |
| First | Cache miss (cold) | Full input token cost |
| Second+ | Cache hit (warm) | 50-90% cheaper (only cache-read tokens billed) |
| Provider | Cache discount |
| Google Gemini | ~90% on cached input tokens |
| Anthropic | ~90% on cached input tokens |
| OpenAI | ~50% on cached input tokens |
12345# Warm the cache mikro cache --context ./docs/ # Now run batch — all questions hit warm cache mikro batch questions.txt --context ./docs/
1mikro cache --context ./docs/ --estimate
1234567891011mikro cache estimate --- context: ./docs/ metadata: Context is a list of 23 items with 144700 total characters, chunk lengths: [5120, 11873, 2960, ...] estimated tokens: 43,500 provider limit: 1,000,000 tokens utilization: 4.3% provider: google model: gemini-3.1-flash-lite-preview ttl: 3600s estimated cost: $0.0033
estimated cost is the input cost of sending the context once, uncached. mikro prints no cached figure; with Gemini's ~90% cache discount, each later question costs about a tenth of it.1mikro batch questions.txt --context ./src/ --max-cost 1.00
| Flag | Default | Description |
--context <path> | — | Context directory or file |
--max-iterations <n> | 30 | Max RLM iterations per question |
--max-cost <n> | — | Max total USD spend |
--parallel <n> | 1 | Accepted; questions currently run one at a time |
--batch-api | false | Accepted; mikro does not submit Gemini Batch API jobs yet |
--verbose | false | Show progress |
--output has no effect here. The CLI reference lists every batch flag.mikro batch accepts --batch-api, but mikro does not submit Gemini Batch API jobs yet. A run with the flag makes the same standard calls as a run without it, and no batch discount applies.| Mode | Input cost (per 1M tokens) | Savings |
| Base (flash-lite) | $0.075 | — |
| + Context caching | ~$0.0075 | 90% |
123456789101112# Prepare questions cat > study.txt << 'EOF' What are the core abstractions? How does the plugin system work? What are the extension points? How is state managed? What patterns does the codebase use? EOF # Warm cache, then batch mikro cache --context ./docs/ mikro batch study.txt --context ./docs/ --max-iterations 5
12345678910cat > audit.txt << 'EOF' Are there any hardcoded credentials? What input validation exists? How are SQL queries constructed? Are there any command injection risks? How is authentication implemented? What error information is leaked to users? EOF mikro batch audit.txt --context ./src/ --ext .ts,.js --max-cost 2.00
12345678910cat > onboard.txt << 'EOF' What is the project structure and architecture? What are the main entry points? How is the database accessed? What external services are called? How are tests organized? What CI/CD pipeline is used? EOF mikro batch onboard.txt --context . --ext .ts,.js,.json,.yaml --tools standard
| Approach | When to use |
| Batch + cache (CAG) | Context fits in provider window, many questions, cost matters |
| Batch + RLM (no cache) | Context too large for system prompt, complex navigation needed |
| Provider | Max context (cached) |
| Google Gemini | 1,000,000 tokens |
| Anthropic | 200,000 tokens |
| OpenAI | 128,000 tokens |
| Amazon Bedrock | 128,000 tokens |