Models
Autohand models
Fantail and Moa are Autohand's coding models. Weka is a typed decision model for release gates, operations, and agent workflows.
Choose a model
| Model | Model ID | Use it for | Context | Max output | Plans |
|---|---|---|---|---|---|
| Fantail | fantail |
Completions, quick fixes, review comments, test scaffolds, and short agent loops. | 256K input | 16K | Free and above |
| Moa | moa |
Large refactors, architecture work, security reviews, and multi-file planning. | 1M input | 262K | Pro and above |
| Weka | weka |
Typed release, risk, triage, and workflow decisions with probabilities. | 32K input | Typed answer | Pro and above |
Fantail and Moa accept provider-qualified aliases as well as their bare identifiers. Weka uses the exact model ID weka at POST /v1/decisions. It is in Public Preview and available only through the Autohand API, Autohand SDK, and Console Playground. The CLI lists Weka for discovery but cannot run it.
Recommended coding models
Use these models for coding agents, long-horizon repo work, tool-heavy sessions, and hosted team workflows.
| Model | Provider | Best for | Context | License | Top signal |
|---|---|---|---|---|---|
| Autohand Fantail | Autohand | Fast coding loops, completions, reviews, and short agent tasks. | 256K | Hosted | Latency-first Autohand model |
| Autohand Moa | Autohand | Larger refactors, planning, and multi-step repository changes. | 1M | Hosted | Reasoning-first Autohand model |
| GLM-5.1 | Z.ai | Long-horizon agentic engineering. | 200K | MIT | Terminal-Bench 2.0 SOTA |
| GLM-5.2 | Z.ai | 1M-context long-horizon agentic engineering. | 1M | MIT | Terminal-Bench 2.1 81.0 |
| Sakana AI Fugu | Sakana AI | Balanced low-latency multi-agent coding and chat. | 1M | Orchestrated | SWE-Bench Pro 59.0 |
| Sakana AI Fugu Ultra | Sakana AI | Multi-agent orchestration across frontier models. | 1M | Orchestrated | SWE-Bench Pro 73.7 |
| Kimi K2.6 | Moonshot AI | Agent swarms and long autonomous runs. | 256K | Modified MIT | SWE-Bench Pro 58.6% |
| DeepSeek V4 Flash | DeepSeek | Quick edits and short loops that do not need the Pro tier. | 1M | MIT | DeepSeek's recommended V4 default |
| DeepSeek V4 Pro | DeepSeek | 1M-context coding and competitive programming. | 1M | MIT | LiveCodeBench 93.5 |
| Qwen3-Coder-Next | Alibaba Qwen | Efficiency per active parameter. | 256K | Apache 2.0 | SWE-Bench Verified 71.3 |
| Qwen3.8-Max | Alibaba Qwen | Long-horizon agentic coding on the Qwen flagship. | 1M | Hosted | 131K max output |
| Qwen3.8-Max-0902 | Alibaba Qwen | Pinning Qwen3.8-Max to a dated snapshot for repeatable runs. | 1M | Hosted | Dated Qwen3.8-Max snapshot |
| Qwen3.8-27B | Alibaba Qwen | Dense repo-level coding. | 1M | Open weights | 131K max output |
| MiniMax M2.5 | MiniMax | Free hosted open-weight coding. | 200K | Open weights | Free hosted variant |
| GLM-5 | Z.ai | Local self-hosting foundations. | 200K | MIT | #1 Vending Bench 2 OSS |
| Devstral 2 | Mistral AI | Mistral coding workloads. | 256K | Apache 2.0 | SWE-Bench Verified 72.2 |
| Devstral Small 2 | Mistral AI | Consumer GPU coding workflows. | 128K | Apache 2.0 | SWE-Bench Verified 68% |
| Trinity Large Thinking | Arcee AI | US-origin open reasoning and long agent loops. | 128K | Apache 2.0 | #2 PinchBench |
Use a model
Pass the model identifier where you start the run. Use fantail for latency-sensitive work and moa when the task needs broad repository context or deliberate reasoning.
CLI
autohand --model fantail --prompt "Find the smallest safe fix for the failing test and apply it." --patch
# Choose Moa for repository-level planning
autohand --model moa --prompt "Map the authentication flow and propose the safest refactor plan." --patch
cURL
curl https://api.autohand.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $AUTOHAND_API_KEY" \
-d '{
"model": "moa",
"messages": [
{"role": "user", "content": "Plan a safe migration for this service."}
]
}'Fantail
Fantail is the latency-first cloud model and the default selection for Autohand Code. Its 256K input context and tool calling support suit focused implementation loops.
- Best for autocomplete, quick bug fixes, code review notes, documentation snippets, and small tests.
- Works well when the relevant files are already in the prompt or easy for the agent to inspect.
- Use Moa instead when the task depends on broad architecture or a large amount of repository context.
Moa
Moa is the reasoning-first cloud model for deeper software engineering work. Its 1M input context, tool calling support, and medium, high, and xhigh reasoning settings fit complex repository work.
- Best for large refactors, migration plans, security audits, technical design, and repository analysis.
- Use it when the model needs to compare patterns across a codebase before making changes.
- Switch back to Fantail for final polish, small follow-up edits, and fast review cycles.
Weka
Weka is trained on open source foundations, distilled from Fantail, and runs through Autohand inference as its first System One Model, trained with Reinforcement Learning for Calibrated Decisions (RLCD).
- Best for release gates, incident triage, workflow routing, and agent checkpoints.
- Use it when application code needs a validated value instead of generated prose.
- Weka is in Public Preview and available only through the API, SDK, and Console Playground.
- Weka returns typed
noul,choice, andscoreanswers with probabilities. - Use
POST /v1/decisionswith model IDweka; Weka does not use the chat-completions contract.
Common patterns
- Start fast. Use Fantail for the first pass when the task is scoped to a file, function, or failing test.
- Escalate for context. Switch to Moa when the answer depends on architecture, cross-package contracts, or hidden coupling.
- Make the branch explicit. Use Weka when the next step needs a typed choice, probability, or score.
- Keep the model explicit. Include the model in scripts and SDK calls so runs are repeatable.