Agent-readable docs index: /llms.txt. Full docs in one file: /llms-full.txt. Download /docs.zip to grep all markdown files locally.

CLI Reference

mikro "query"

Run an RLM query against a context.
mikro "query" [options]
The default command. Loads context into a Python REPL, then iterates with the LLM until it produces a final answer.

Options

FlagTypeDefaultDescription
--context <path>string—Path to context directory or file
--output <mode>stringtextOutput mode: text, json, or stream
--verbosebooleanfalseShow iteration progress on stderr
--max-iterations <n>number30Maximum RLM iterations before forced termination
--timeout <ms>number300000Timeout in milliseconds (5 minutes)
--statsbooleanfalseEmit JSON stats to stderr (or include in --output json)
--log <path>string—Write structured JSONL log to file
--tools <level>stringcoreTool level: core, standard, or full
--max-cost <n>number—Maximum USD spend per run
--max-tokens <n>number—Maximum total tokens per run
--max-depth <n>number—Maximum recursive rlm_query depth
--model <ref>stringfrom configModel for this run: provider/model, or a bare model id on the configured provider. Outranks settings.json and mikro.yaml
--ext <list>string.mdFile extensions for context dirs (comma-separated)
--thinking <level>string—Thinking level: minimal, low, medium, high. Sent to any provider whose model supports a reasoning level
--temperature <n>numberunsetSampling temperature, 0 to 2. Unset sends no temperature, so the provider decides
--cachebooleanfalseEnable CAG mode (full context cached in system prompt)
--no-sessionbooleanfalseSkip saving the run under ~/.mikro/sessions/

Context loading

InputBehavior
--context dir/Recursively reads files matching --ext as list[{path, content}]
--context file.mdReads as single string
--context file.jsonParses JSON as dict or list
stdin pipeRead as the query when no query argument is given. With a query argument, stdin is ignored

Examples

# Basic query with directory context mikro "How does IPC work?" --context ./docs/ # JSON output with stats mikro "Summarize this" --context paper.md --output json --stats # Code analysis with extended file types mikro "Analyze code" --context ./src/ --tools full --ext .ts,.js # Budget-limited query mikro "Quick question" --max-cost 0.10 --max-tokens 5000 # Query read from stdin, with logging echo "Summarize this" | mikro --context paper.md --log run.jsonl # CAG mode for repeated queries mikro "First question" --context ./docs/ --cache mikro "Follow-up question" --context ./docs/ --cache # Higher thinking level mikro "Complex analysis" --context ./src/ --thinking high

Output formats

{ "answer": "The answer to your query...", "references": ["docs/start/create-project.md", "docs/concept/ipc.md"], "usage": { "inputTokens": 12500, "outputTokens": 3200, "cacheReadTokens": 0, "cacheWriteTokens": 0, "totalCost": 0.0041, "llmCalls": 5 }, "iterations": 3, "model": "google/gemini-3.1-flash-lite-preview" }
--output json prints the object on one line. Besides the fields above it can carry budgetHit, usageBreakdown, geminiCounts, geminiBatteriesUsed, and stats (with --stats), and usage can carry reasoningTokens. Stream iterations count from 0, and the final event's iterations is the total.

mikro init

Scaffold a .mikro/ config directory with templates.
mikro init [--template <type>] [--dir <path>]
FlagTypeDefaultDescription
--template <type>stringdefaultTemplate type: default or code
--dir <path>string. (cwd)Directory to scaffold in
Creates a .mikro/ directory containing:
FilePurpose
mikro.yamlMain configuration (model, budget, context, storage)
SYSTEM.mdSystem prompt used by the RLM loop
CRITERIA.mdOutput criteria for quality checks
TOOLS.mdCustom Python tools exposed to the RLM
mikro reads configuration only from the .mikro/ directory. A mikro.yaml or SYSTEM.md outside .mikro/ is ignored. For a project set up before the rename to mikro, see Configuration.

Templates

  • default: General-purpose RLM usage with balanced system prompt and criteria
  • code: Code analysis template with a code-focused system prompt and criteria, code file extensions in context.extensions, and tools-level: standard

Example

# Scaffold with default template mikro init # Scaffold with code analysis template mikro init --template code # Scaffold in a specific directory mikro init --template default --dir ./my-project

mikro cache

Pre-warm the provider cache for a given context, or estimate its size and cost without making any API calls.
mikro cache --context <path> [--estimate] [options]
--context is required. Without --estimate, mikro performs a single-iteration warmup run so the provider caches the prompt prefix for subsequent queries.
FlagTypeDefaultDescription
--context <path>stringrequiredPath to context directory or file
--estimatebooleanfalsePrint token/cost estimate only — skip the warmup call
--ext <list>stringfrom mikro.yamlFile extensions when --context is a directory (comma-separated)
--tools <level>stringcoreTool level: core, standard, or full
--timeout <ms>number300000Warmup timeout
--verbosebooleanfalseVerbose stderr logging

Examples

# Estimate only, no LLM calls mikro cache --context ./docs/ --estimate # Warm the provider cache for future queries mikro cache --context ./docs/ # Custom extensions for a source-code corpus mikro cache --context ./src/ --ext .ts,.js --estimate

Outputs

With --estimate, a key: value block is printed to stdout and the process exits without calling any LLM:
mikro cache estimate --- context: ./docs/ metadata: Context is a list of 23 items with 144700 total characters, chunk lengths: [5120, 11873, 2960, ...] estimated tokens: 43,500 provider limit: 1,000,000 tokens utilization: 4.3% provider: google model: gemini-3.1-flash-lite-preview ttl: 3600s estimated cost: $0.0033
Without --estimate, a minimal RLM loop runs to prime the provider cache. Progress and the summary are emitted to stderr:
mikro: warming cache for ./docs/ (~43,500 tokens) mikro: cache warmup complete provider: google model: gemini-3.1-flash-lite-preview estimated tokens: 43,500 ttl: 3600s estimated cost: $0.0033
If the context exceeds the provider's token limit, the command exits with a non-zero status and an error message. See Cache Mode for the full caching workflow.

mikro batch

Run bulk queries from a questions file against a shared cached context. Each question is executed through the same RLM loop used by mikro "query", with provider-level prompt caching always enabled so the first question pays full price and subsequent questions benefit from the cache.
mikro batch <questions-file> [options]
Questions are read one per line. Blank lines and lines beginning with # are ignored.
FlagTypeDefaultDescription
<questions-file>pathrequiredPath to a text file of questions (one per line)
--context <path>string—Shared context for every question
--max-iterations <n>number30Maximum RLM iterations per question
--timeout <ms>number300000Per-question timeout
--max-cost <n>number—Stop after cumulative USD cost crosses this threshold
--max-tokens <n>number—Per-question token cap
--max-depth <n>number—Maximum recursive rlm_query depth
--parallel <n>number1Concurrency hint (currently executes sequentially)
--batch-apibooleanfalseAccepted, but mikro does not submit Gemini Batch API jobs yet, so the run is unchanged
--tools <level>stringcoreTool level: core, standard, or full
--ext <list>stringfrom mikro.yamlFile extensions for directory context
--verbosebooleanfalseShow per-question progress on stderr
--cache does not need to be passed: batch mode always runs with cache.enabled = true. If the context exceeds the provider's token limit, mikro falls back to pgserve storage mode when storage.enabled is auto or always.

Example

# Run a questions file against a cached docs corpus mikro batch questions.txt --context ./docs/ # Stop if total spend crosses $1.00 mikro batch questions.txt --context ./src/ --max-cost 1.00

Outputs

mikro batch writes JSONL to stdout: one JSON object per question, followed by a final aggregate line:
{"question":"How does IPC work?","answer":"IPC uses...","stats":{"iterations":2,"inputTokens":42100,"outputTokens":520,"cost":0.0042}} {"question":"Where is auth defined?","answer":"src/auth.ts...","stats":{"iterations":1,"inputTokens":820,"outputTokens":310,"cost":0.0009}} {"type":"aggregate","total_questions":2,"completed":2,"total_cost":0.0051,"cache_savings":0.004}
Budget trips, cache fallbacks, and verbose progress are logged to stderr so the stdout stream stays valid JSONL for downstream pipelines. See Batch Mode for full details.

mikro stats

Query run history and cost breakdowns from the mikro observability database (pgserve, in storage.data-dir, which defaults to ~/.mikro/data). A run records there only when it uses storage mode: storage.enabled: always, or auto when the context exceeds the provider limit.
mikro stats [options]
FlagTypeDefaultDescription
--run <id>string—Show the event timeline for a specific session id
--costsbooleanfalseShow cost breakdown grouped by model
--toolsbooleanfalseShow REPL tool usage grouped by session
--since <duration>string—Limit --costs or --tools to the recent window (30m, 24h, 7d)
--output jsonliteral—Emit structured JSON instead of the terminal table
Without any flags, mikro stats prints the 20 most recent sessions as a terminal table (id, query, model, iterations, cost, status, duration).
Stats require pgserve storage. If the data directory (~/.mikro/data by default) does not exist, the command prints "No stats yet. Run a query first." and exits cleanly. See Configuration for storage setup.

Examples

# Most recent 20 runs as a table mikro stats # JSON for scripting / jq pipelines mikro stats --output json # Cost by model over the last 24 hours mikro stats --costs --since 24h # Tool usage over the last week mikro stats --tools --since 7d # Full event timeline for a specific run mikro stats --run 0c3e2f1a-...-9f02

Outputs

Default (sessions table) — plain-text columns written to stdout:
ID Query Model Iter Cost Status Duration -------------------------------------------------------------------------------------------------------------- 0c3e2f1a.. How does IPC work? google/gemini-3.1-flash... 3 $0.0042 completed 4.1s f91d8a05.. Summarize paper.md google/gemini-3.1-flash... 2 $0.0011 completed 1.8s
--run <id> — one row per event (llm_call, repl_exec, sub_call) with iteration, token counts, cost, duration, and kind-specific detail (model, code preview, request type).
--costs — one row per (session, model) pair with total calls, input/output tokens, cost, and average call duration.
--tools — one row per (session, request_type) with calls, errors, and average duration.
--output json — any of the above as a pretty-printed JSON array of rows.

mikro benchmark

Run benchmarks that compare the RLM loop against a direct LLM call on the same question. mikro benchmark does not accept --context: each mode ships its own dataset.
mikro benchmark <mode> [options]
ModeDatasetMeasures
costBuilt-in curated dataset (src/benchmark-data.json)Tokens, cost, latency savings for RLM vs direct
oolongOolong Synth auto-downloaded via HuggingFaceAnswer quality (accuracy) plus the same cost metrics

Flags

FlagApplies toDefaultDescription
--output jsoncosttablePrint JSON results to stdout instead of the table
--samples <n>oolong5Number of samples to evaluate
--idx <n>oolong—Run a specific sample index only (ignores --samples)
--tools <level>bothcoreTool level used for the RLM runs
Model and provider are resolved from .mikro/mikro.yaml and ~/.mikro/settings.json in the usual priority order.

Examples

# Cost benchmark, formatted table on stderr mikro benchmark cost # Cost benchmark, machine-readable JSON on stdout mikro benchmark cost --output json # Oolong quality run, 5 samples (default) mikro benchmark oolong # Oolong, 20 samples mikro benchmark oolong --samples 20 # Oolong, sample index 42 only mikro benchmark oolong --idx 42

Outputs

Both modes print a box-drawn comparison table to stderr with per-question rows (Direct / RLM / Savings) and a TOTALS footer covering tokens, cost, latency, and average RLM iterations:
┌───────────────┬──────────┬─────────┬──────────┬──────────┬──────┐ │ Question │ Mode │ Tokens │ Cost │ Latency │ Iter │ ├───────────────┼──────────┼─────────┼──────────┼──────────┼──────┤ │ ipc_summary │ Direct │ 12,400 │ $0.00093 │ 820ms│ - │ │ │ RLM │ 3,100 │ $0.00024 │ 2,410ms│ 3 │ │ │ Savings │ 75.0% │ 74.2% │ - │ │ ├───────────────┼──────────┼─────────┼──────────┼──────────┼──────┤ │ TOTALS │ Direct │ 94,200 │ $0.00707 │ 6.2s │ - │ │ │ RLM │ 23,500 │ $0.00182 │ 18.4s │ 2.8 │ │ │ Savings │ 75.1% │ 74.3% │ - │ │ └───────────────┴──────────┴─────────┴──────────┴──────────┴──────┘
With mikro benchmark cost --output json, the same results are emitted as a structured JSON document to stdout (timestamp, mode, model, per-question runs[], and totals). Every benchmark, table or JSON, is also persisted to ~/.mikro/benchmarks/benchmark-<mode>-<timestamp>.json and the saved path is printed to stderr.

mikro config

Manage global settings stored at ~/.mikro/settings.json.

mikro config set

mikro config set <key> <value>
Set a configuration value. Values are type-coerced: "true" becomes boolean, numeric strings become numbers.
mikro config set GEMINI_API_KEY sk-abc123 mikro config set model.provider google mikro config set model.model gemini-3.1-flash-lite-preview

mikro config get

mikro config get <key>
Retrieve a setting value. API keys are masked in output.
$ mikro config get model.provider google

mikro config list

mikro config list
Show all configured settings. Sensitive keys (containing API_KEY, SECRET, TOKEN) are masked.

mikro config delete

mikro config delete <key>
Remove a setting.

mikro config path

mikro config path
Print the settings file path (~/.mikro/settings.json).

Common keys

KeyDescriptionExample
GEMINI_API_KEYGoogle Gemini API keyAIza...
ANTHROPIC_API_KEYAnthropic API keysk-ant-...
OPENAI_API_KEYOpenAI API keysk-...
GROQ_API_KEYGroq API keygsk_...
XAI_API_KEYxAI API keyxai-...
OPENROUTER_API_KEYOpenRouter API keysk-or-...
model.providerLLM providergoogle, anthropic, openai
model.modelModel IDgemini-3.1-flash-lite-preview
model.sub-call-modelModel for llm_query() sub-callsgemini-3.1-flash-lite-preview
mikro also reads GOOGLE_API_KEY, DEEPSEEK_API_KEY, KIMI_API_KEY, MOONSHOT_API_KEY, MINIMAX_API_KEY, ZAI_API_KEY and GLM_API_KEY from settings. A key stored here is used only when the matching environment variable is unset.
mikro config set stores any key you give it, and the help that mikro config prints also lists model.sub_call_model, budget.*, tools_level, cache.retention and gemini.* keys. mikro reads none of those from settings.json (for the sub-call model it reads only the hyphenated model.sub-call-model): set budgets, tool level, cache and Gemini options in .mikro/mikro.yaml or with a flag.

Priority order

Settings are resolved in this order (highest priority first):
  1. CLI flags (--max-cost 0.10, --model openai/gpt-4o)
  2. Global ~/.mikro/settings.json, for model.provider, model.model and model.sub-call-model
  3. Project .mikro/mikro.yaml
  4. Hardcoded defaults

Other commands

CommandWhat it does
mikro doctorPrints the version, which of six provider keys are set (Gemini, OpenAI, Anthropic, Groq, xAI, OpenRouter), config-declared providers, whether the configured model resolves, and the settings file. Exits 1 when any of those six keys is unset
mikro update [--force]Fetches the newest main commit of the git install and rebuilds it in place. Refuses a checkout with local changes unless --force is given
mikro migrate [--apply]Finds config directories, .mcp.json entries and Claude Code plugin registrations left from before the rename to mikro, and rewrites them. Dry run unless --apply; see Configuration
mikro mcp [--dir <path>]Runs a stdio MCP server that exposes mikro_query plus one tool per microagent; --dir sets the project root
mikro acpRuns a stdio Agent Client Protocol agent (experimental)
To give Claude Code the MCP tools:
claude mcp add mikro -- mikro mcp

Tool levels

The --tools flag controls which functions are available in the REPL:

core (default)

Paper-faithful RLM functions:
FunctionDescription
contextInjected context variable
llm_query(prompt)Single LLM completion
llm_query_batched(prompts)Concurrent LLM calls
rlm_query(prompt)Recursive child RLM session
rlm_query_batched(prompts)Parallel child RLM sessions
SHOW_VARS()List all REPL variables
FINAL(answer)Terminate with answer string
FINAL_VAR(name)Terminate with variable value
Plus any custom tools defined in .mikro/TOOLS.md.

standard (core + batteries)

All core functions plus utility batteries:
FunctionDescription
describe_context()Metadata overview of loaded context
preview_context()Content sample
search_context(query)Keyword search over context, ranked by match count
grep_context(pattern)Regex search over context
chunk_context()Split context into chunks
chunk_text(text)Split arbitrary text by size
map_query(items, template)Run one LLM call per item through a template with an {item} placeholder
reduce_query(results, prompt)Combine results into one answer with a prompt holding a {results} placeholder
run_cli(cmd, *args)Run a command and capture its returncode, stdout and stderr
With Google provider, also includes Gemini batteries:
FunctionDescription
web_search(query)Google web search
fetch_url(url)Fetch and summarize URL content
generate_image(prompt)Image generation

full (standard + environment)

All standard functions plus auto-injected information about available Python packages and versions in the REPL environment.

Global flags

These flags work with any command:
FlagDescription
--help, -hShow help message
--version, -vShow version
--schemaPrint the machine-readable CLI schema (flags, JSON output, exit codes) as JSON