mikro "query"1mikro "query" [options]
| Flag | Type | Default | Description |
--context <path> | string | — | Path to context directory or file |
--output <mode> | string | text | Output mode: text, json, or stream |
--verbose | boolean | false | Show iteration progress on stderr |
--max-iterations <n> | number | 30 | Maximum RLM iterations before forced termination |
--timeout <ms> | number | 300000 | Timeout in milliseconds (5 minutes) |
--stats | boolean | false | Emit JSON stats to stderr (or include in --output json) |
--log <path> | string | — | Write structured JSONL log to file |
--tools <level> | string | core | Tool level: core, standard, or full |
--max-cost <n> | number | — | Maximum USD spend per run |
--max-tokens <n> | number | — | Maximum total tokens per run |
--max-depth <n> | number | — | Maximum recursive rlm_query depth |
--model <ref> | string | from config | Model for this run: provider/model, or a bare model id on the configured provider. Outranks settings.json and mikro.yaml |
--ext <list> | string | .md | File extensions for context dirs (comma-separated) |
--thinking <level> | string | — | Thinking level: minimal, low, medium, high. Sent to any provider whose model supports a reasoning level |
--temperature <n> | number | unset | Sampling temperature, 0 to 2. Unset sends no temperature, so the provider decides |
--cache | boolean | false | Enable CAG mode (full context cached in system prompt) |
--no-session | boolean | false | Skip saving the run under ~/.mikro/sessions/ |
| Input | Behavior |
--context dir/ | Recursively reads files matching --ext as list[{path, content}] |
--context file.md | Reads as single string |
--context file.json | Parses JSON as dict or list |
| stdin pipe | Read as the query when no query argument is given. With a query argument, stdin is ignored |
123456789101112131415161718192021# Basic query with directory context mikro "How does IPC work?" --context ./docs/ # JSON output with stats mikro "Summarize this" --context paper.md --output json --stats # Code analysis with extended file types mikro "Analyze code" --context ./src/ --tools full --ext .ts,.js # Budget-limited query mikro "Quick question" --max-cost 0.10 --max-tokens 5000 # Query read from stdin, with logging echo "Summarize this" | mikro --context paper.md --log run.jsonl # CAG mode for repeated queries mikro "First question" --context ./docs/ --cache mikro "Follow-up question" --context ./docs/ --cache # Higher thinking level mikro "Complex analysis" --context ./src/ --thinking high
1234567891011121314{ "answer": "The answer to your query...", "references": ["docs/start/create-project.md", "docs/concept/ipc.md"], "usage": { "inputTokens": 12500, "outputTokens": 3200, "cacheReadTokens": 0, "cacheWriteTokens": 0, "totalCost": 0.0041, "llmCalls": 5 }, "iterations": 3, "model": "google/gemini-3.1-flash-lite-preview" }
--output json prints the object on one line. Besides the fields above it can carry budgetHit, usageBreakdown, geminiCounts, geminiBatteriesUsed, and stats (with --stats), and usage can carry reasoningTokens. Stream iterations count from 0, and the final event's iterations is the total.mikro init.mikro/ config directory with templates.1mikro init [--template <type>] [--dir <path>]
| Flag | Type | Default | Description |
--template <type> | string | default | Template type: default or code |
--dir <path> | string | . (cwd) | Directory to scaffold in |
.mikro/ directory containing:| File | Purpose |
mikro.yaml | Main configuration (model, budget, context, storage) |
SYSTEM.md | System prompt used by the RLM loop |
CRITERIA.md | Output criteria for quality checks |
TOOLS.md | Custom Python tools exposed to the RLM |
.mikro/ directory. A mikro.yaml or SYSTEM.md outside .mikro/ is ignored. For a project set up before the rename to mikro, see Configuration.default: General-purpose RLM usage with balanced system prompt and criteriacode: Code analysis template with a code-focused system prompt and criteria, code file extensions in context.extensions, and tools-level: standard12345678# Scaffold with default template mikro init # Scaffold with code analysis template mikro init --template code # Scaffold in a specific directory mikro init --template default --dir ./my-project
mikro cache1mikro cache --context <path> [--estimate] [options]
--context is required. Without --estimate, mikro performs a single-iteration warmup run so the provider caches the prompt prefix for subsequent queries.| Flag | Type | Default | Description |
--context <path> | string | required | Path to context directory or file |
--estimate | boolean | false | Print token/cost estimate only — skip the warmup call |
--ext <list> | string | from mikro.yaml | File extensions when --context is a directory (comma-separated) |
--tools <level> | string | core | Tool level: core, standard, or full |
--timeout <ms> | number | 300000 | Warmup timeout |
--verbose | boolean | false | Verbose stderr logging |
12345678# Estimate only, no LLM calls mikro cache --context ./docs/ --estimate # Warm the provider cache for future queries mikro cache --context ./docs/ # Custom extensions for a source-code corpus mikro cache --context ./src/ --ext .ts,.js --estimate
--estimate, a key: value block is printed to stdout and the process exits without calling any LLM:1234567891011mikro cache estimate --- context: ./docs/ metadata: Context is a list of 23 items with 144700 total characters, chunk lengths: [5120, 11873, 2960, ...] estimated tokens: 43,500 provider limit: 1,000,000 tokens utilization: 4.3% provider: google model: gemini-3.1-flash-lite-preview ttl: 3600s estimated cost: $0.0033
--estimate, a minimal RLM loop runs to prime the provider cache. Progress and the summary are emitted to stderr:1234567mikro: warming cache for ./docs/ (~43,500 tokens) mikro: cache warmup complete provider: google model: gemini-3.1-flash-lite-preview estimated tokens: 43,500 ttl: 3600s estimated cost: $0.0033
mikro batchmikro "query", with provider-level prompt caching always enabled so the first question pays full price and subsequent questions benefit from the cache.1mikro batch <questions-file> [options]
# are ignored.| Flag | Type | Default | Description |
<questions-file> | path | required | Path to a text file of questions (one per line) |
--context <path> | string | — | Shared context for every question |
--max-iterations <n> | number | 30 | Maximum RLM iterations per question |
--timeout <ms> | number | 300000 | Per-question timeout |
--max-cost <n> | number | — | Stop after cumulative USD cost crosses this threshold |
--max-tokens <n> | number | — | Per-question token cap |
--max-depth <n> | number | — | Maximum recursive rlm_query depth |
--parallel <n> | number | 1 | Concurrency hint (currently executes sequentially) |
--batch-api | boolean | false | Accepted, but mikro does not submit Gemini Batch API jobs yet, so the run is unchanged |
--tools <level> | string | core | Tool level: core, standard, or full |
--ext <list> | string | from mikro.yaml | File extensions for directory context |
--verbose | boolean | false | Show per-question progress on stderr |
--cache does not need to be passed: batch mode always runs with cache.enabled = true. If the context exceeds the provider's token limit, mikro falls back to pgserve storage mode when storage.enabled is auto or always.12345# Run a questions file against a cached docs corpus mikro batch questions.txt --context ./docs/ # Stop if total spend crosses $1.00 mikro batch questions.txt --context ./src/ --max-cost 1.00
mikro batch writes JSONL to stdout: one JSON object per question, followed by a final aggregate line:123{"question":"How does IPC work?","answer":"IPC uses...","stats":{"iterations":2,"inputTokens":42100,"outputTokens":520,"cost":0.0042}} {"question":"Where is auth defined?","answer":"src/auth.ts...","stats":{"iterations":1,"inputTokens":820,"outputTokens":310,"cost":0.0009}} {"type":"aggregate","total_questions":2,"completed":2,"total_cost":0.0051,"cache_savings":0.004}
mikro statsstorage.data-dir, which defaults to ~/.mikro/data). A run records there only when it uses storage mode: storage.enabled: always, or auto when the context exceeds the provider limit.1mikro stats [options]
| Flag | Type | Default | Description |
--run <id> | string | — | Show the event timeline for a specific session id |
--costs | boolean | false | Show cost breakdown grouped by model |
--tools | boolean | false | Show REPL tool usage grouped by session |
--since <duration> | string | — | Limit --costs or --tools to the recent window (30m, 24h, 7d) |
--output json | literal | — | Emit structured JSON instead of the terminal table |
mikro stats prints the 20 most recent sessions as a terminal table (id, query, model, iterations, cost, status, duration).~/.mikro/data by default) does not exist, the command prints "No stats yet. Run a query first." and exits cleanly. See Configuration for storage setup.1234567891011121314# Most recent 20 runs as a table mikro stats # JSON for scripting / jq pipelines mikro stats --output json # Cost by model over the last 24 hours mikro stats --costs --since 24h # Tool usage over the last week mikro stats --tools --since 7d # Full event timeline for a specific run mikro stats --run 0c3e2f1a-...-9f02
1234ID Query Model Iter Cost Status Duration -------------------------------------------------------------------------------------------------------------- 0c3e2f1a.. How does IPC work? google/gemini-3.1-flash... 3 $0.0042 completed 4.1s f91d8a05.. Summarize paper.md google/gemini-3.1-flash... 2 $0.0011 completed 1.8s
--run <id> — one row per event (llm_call, repl_exec, sub_call) with iteration, token counts, cost, duration, and kind-specific detail (model, code preview, request type).--costs — one row per (session, model) pair with total calls, input/output tokens, cost, and average call duration.--tools — one row per (session, request_type) with calls, errors, and average duration.--output json — any of the above as a pretty-printed JSON array of rows.mikro benchmarkmikro benchmark does not accept --context: each mode ships its own dataset.1mikro benchmark <mode> [options]
| Mode | Dataset | Measures |
cost | Built-in curated dataset (src/benchmark-data.json) | Tokens, cost, latency savings for RLM vs direct |
oolong | Oolong Synth auto-downloaded via HuggingFace | Answer quality (accuracy) plus the same cost metrics |
| Flag | Applies to | Default | Description |
--output json | cost | table | Print JSON results to stdout instead of the table |
--samples <n> | oolong | 5 | Number of samples to evaluate |
--idx <n> | oolong | — | Run a specific sample index only (ignores --samples) |
--tools <level> | both | core | Tool level used for the RLM runs |
.mikro/mikro.yaml and ~/.mikro/settings.json in the usual priority order.1234567891011121314# Cost benchmark, formatted table on stderr mikro benchmark cost # Cost benchmark, machine-readable JSON on stdout mikro benchmark cost --output json # Oolong quality run, 5 samples (default) mikro benchmark oolong # Oolong, 20 samples mikro benchmark oolong --samples 20 # Oolong, sample index 42 only mikro benchmark oolong --idx 42
1234567891011┌───────────────┬──────────┬─────────┬──────────┬──────────┬──────┐ │ Question │ Mode │ Tokens │ Cost │ Latency │ Iter │ ├───────────────┼──────────┼─────────┼──────────┼──────────┼──────┤ │ ipc_summary │ Direct │ 12,400 │ $0.00093 │ 820ms│ - │ │ │ RLM │ 3,100 │ $0.00024 │ 2,410ms│ 3 │ │ │ Savings │ 75.0% │ 74.2% │ - │ │ ├───────────────┼──────────┼─────────┼──────────┼──────────┼──────┤ │ TOTALS │ Direct │ 94,200 │ $0.00707 │ 6.2s │ - │ │ │ RLM │ 23,500 │ $0.00182 │ 18.4s │ 2.8 │ │ │ Savings │ 75.1% │ 74.3% │ - │ │ └───────────────┴──────────┴─────────┴──────────┴──────────┴──────┘
mikro benchmark cost --output json, the same results are emitted as a structured JSON document to stdout (timestamp, mode, model, per-question runs[], and totals). Every benchmark, table or JSON, is also persisted to ~/.mikro/benchmarks/benchmark-<mode>-<timestamp>.json and the saved path is printed to stderr.mikro config~/.mikro/settings.json.mikro config set1mikro config set <key> <value>
"true" becomes boolean, numeric strings become numbers.123mikro config set GEMINI_API_KEY sk-abc123 mikro config set model.provider google mikro config set model.model gemini-3.1-flash-lite-preview
mikro config get1mikro config get <key>
12$ mikro config get model.provider google
mikro config list1mikro config list
API_KEY, SECRET, TOKEN) are masked.mikro config delete1mikro config delete <key>
mikro config path1mikro config path
~/.mikro/settings.json).| Key | Description | Example |
GEMINI_API_KEY | Google Gemini API key | AIza... |
ANTHROPIC_API_KEY | Anthropic API key | sk-ant-... |
OPENAI_API_KEY | OpenAI API key | sk-... |
GROQ_API_KEY | Groq API key | gsk_... |
XAI_API_KEY | xAI API key | xai-... |
OPENROUTER_API_KEY | OpenRouter API key | sk-or-... |
model.provider | LLM provider | google, anthropic, openai |
model.model | Model ID | gemini-3.1-flash-lite-preview |
model.sub-call-model | Model for llm_query() sub-calls | gemini-3.1-flash-lite-preview |
GOOGLE_API_KEY, DEEPSEEK_API_KEY, KIMI_API_KEY, MOONSHOT_API_KEY, MINIMAX_API_KEY, ZAI_API_KEY and GLM_API_KEY from settings. A key stored here is used only when the matching environment variable is unset.mikro config set stores any key you give it, and the help that mikro config prints also lists model.sub_call_model, budget.*, tools_level, cache.retention and gemini.* keys. mikro reads none of those from settings.json (for the sub-call model it reads only the hyphenated model.sub-call-model): set budgets, tool level, cache and Gemini options in .mikro/mikro.yaml or with a flag.--max-cost 0.10, --model openai/gpt-4o)~/.mikro/settings.json, for model.provider, model.model and model.sub-call-model.mikro/mikro.yaml| Command | What it does |
mikro doctor | Prints the version, which of six provider keys are set (Gemini, OpenAI, Anthropic, Groq, xAI, OpenRouter), config-declared providers, whether the configured model resolves, and the settings file. Exits 1 when any of those six keys is unset |
mikro update [--force] | Fetches the newest main commit of the git install and rebuilds it in place. Refuses a checkout with local changes unless --force is given |
mikro migrate [--apply] | Finds config directories, .mcp.json entries and Claude Code plugin registrations left from before the rename to mikro, and rewrites them. Dry run unless --apply; see Configuration |
mikro mcp [--dir <path>] | Runs a stdio MCP server that exposes mikro_query plus one tool per microagent; --dir sets the project root |
mikro acp | Runs a stdio Agent Client Protocol agent (experimental) |
1claude mcp add mikro -- mikro mcp
--tools flag controls which functions are available in the REPL:core (default)| Function | Description |
context | Injected context variable |
llm_query(prompt) | Single LLM completion |
llm_query_batched(prompts) | Concurrent LLM calls |
rlm_query(prompt) | Recursive child RLM session |
rlm_query_batched(prompts) | Parallel child RLM sessions |
SHOW_VARS() | List all REPL variables |
FINAL(answer) | Terminate with answer string |
FINAL_VAR(name) | Terminate with variable value |
.mikro/TOOLS.md.standard (core + batteries)| Function | Description |
describe_context() | Metadata overview of loaded context |
preview_context() | Content sample |
search_context(query) | Keyword search over context, ranked by match count |
grep_context(pattern) | Regex search over context |
chunk_context() | Split context into chunks |
chunk_text(text) | Split arbitrary text by size |
map_query(items, template) | Run one LLM call per item through a template with an {item} placeholder |
reduce_query(results, prompt) | Combine results into one answer with a prompt holding a {results} placeholder |
run_cli(cmd, *args) | Run a command and capture its returncode, stdout and stderr |
| Function | Description |
web_search(query) | Google web search |
fetch_url(url) | Fetch and summarize URL content |
generate_image(prompt) | Image generation |
full (standard + environment)| Flag | Description |
--help, -h | Show help message |
--version, -v | Show version |
--schema | Print the machine-readable CLI schema (flags, JSON output, exit codes) as JSON |