Availability

Weka is in Public Preview on Pro, Max, Team, and Enterprise. It is available only through the Autohand API, Autohand SDK, and Console Playground. Use the model ID weka with POST /v1/decisions.

CLI support: The Autohand CLI lists Weka for discovery, but cannot run it. Weka uses a structured decision contract rather than chat completions.

Weka is trained on open source foundations, distilled from Fantail, and runs through Autohand inference as its first System One Model, trained with Reinforcement Learning for Calibrated Decisions (RLCD).

Weka uses a structured decision contract. It does not accept chat messages and cannot be called through /v1/chat/completions.

API and SDK calls use the same structured state, questions, answers, and quota checks.

The request contract

A Weka request has three top-level fields:

FieldMeaning
modelThe literal value weka.
stateThe JSON facts Weka should evaluate. It can be a string, number, boolean, object, array, or null.
questionsA non-empty object of named noul, choice, or score questions.

Autohand validates at least one named question before metering. A choice question needs at least one named criterion, and a score question needs at least two ordered criteria.

Question types

TypeUse it forAnswer
noulA yes or no question where uncertainty matters.A noul probability from 0 to 1.
choiceSelecting one named option.The selected key, confidence, and probability per option.
scoreAn ordered scale with at least two criteria.A numeric score, confidence, legend, and distribution.

cURL

curl https://api.autohand.ai/v1/decisions \
  -H "Authorization: Bearer $AUTOHAND_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "weka",
    "state": {
      "tests": "passed",
      "changed_files": 14,
      "touches_payments": true
    },
    "questions": {
      "needs_review": {
        "type": "noul",
        "instructions": "Does this release need human review?"
      },
      "release_lane": {
        "type": "choice",
        "instructions": "Choose the safest release lane.",
        "criteria": {
          "canary": "Deploy to a small cohort first",
          "stable": "Proceed to the normal rollout",
          "blocked": "Do not deploy"
        }
      },
      "risk": {
        "type": "score",
        "instructions": "Score operational risk.",
        "criteria": ["Low", "Medium", "High"]
      }
    }
  }'

TypeScript

type WekaDecision = {
  model: "weka";
  answers: Record<string, unknown>;
  usage: { input_tokens: number; output_tokens: number };
};

const response = await fetch("https://api.autohand.ai/v1/decisions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AUTOHAND_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "weka",
    state: releaseState,
    questions: releaseQuestions,
  }),
});

if (!response.ok) throw new Error(await response.text());
const decision = (await response.json()) as WekaDecision;

Python

import os
import requests

response = requests.post(
    "https://api.autohand.ai/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['AUTOHAND_API_KEY']}"},
    json={
        "model": "weka",
        "state": release_state,
        "questions": release_questions,
    },
    timeout=30,
)
response.raise_for_status()
decision = response.json()

Tutorial 1

Build a release gate

Keep facts and policy separate. Send the current build facts to Weka, then let deterministic application code decide which probabilities are acceptable.

1. Describe observable state

const releaseState = {
  commit: process.env.GIT_SHA,
  tests: { unit: "passed", integration: "passed" },
  changedFiles: 14,
  touchesPayments: true,
  previousIncidents: 1,
};

2. Ask for values your pipeline can use

const releaseQuestions = {
  needs_review: {
    type: "noul",
    instructions: "Does this change require human review before deployment?",
  },
  release_lane: {
    type: "choice",
    instructions: "Choose the safest release lane from the supplied state.",
    criteria: {
      canary: "Release to 5% of traffic first",
      stable: "Use the normal rollout",
      blocked: "Do not deploy",
    },
  },
  risk: {
    type: "score",
    instructions: "Score the operational risk.",
    criteria: ["Low", "Medium", "High", "Critical"],
  },
};

3. Apply a deterministic threshold

const lane = decision.answers.release_lane;
const review = decision.answers.needs_review;

if (lane.type !== "choice" || review.type !== "noul") {
  throw new Error("Unexpected Weka answer shape");
}

if (lane.choice === "blocked" || review.noul >= 0.8) {
  await requestHumanReview(decision);
} else if (lane.choice === "canary" && lane.confidence >= 0.7) {
  await deploy({ trafficPercent: 5 });
} else {
  await holdRelease(decision);
}

Do not treat confidence as permission by itself. Validate the answer type and selected key, then enforce thresholds in code that can be reviewed and tested.

Tutorial 2

Triage an incident

Build one state object from observability data and recent delivery events. Ask Weka for severity and the next runbook branch.

questions = {
    "severity": {
        "type": "score",
        "instructions": "Score customer impact.",
        "criteria": ["No impact", "Degraded", "Major", "Critical"],
    },
    "response": {
        "type": "choice",
        "instructions": "Choose the next runbook branch.",
        "criteria": {
            "rollback": "The latest deployment is the likely cause",
            "investigate": "Evidence is incomplete or points elsewhere",
            "monitor": "The service is recovering without intervention",
        },
    },
}

decision = decide(state=incident_snapshot, questions=questions)
branch = decision["answers"]["response"]

if branch["confidence"] < 0.65:
    page_on_call(decision)
elif branch["choice"] == "rollback":
    run_rollback()

Tutorial 3

Put Weka inside an agent workflow

An agent can use Weka after it has gathered evidence. A low confidence result can send the agent back for more information, while a blocked decision can hand control to a person.

for (let attempt = 0; attempt < 3; attempt += 1) {
  const decision = await decideWithWeka({
    patch: await agent.currentDiff(),
    checks: await collectCheckResults(),
    openQuestions: await agent.openQuestions(),
  });

  const next = decision.answers.next_step;
  if (next.type !== "choice") throw new Error("Invalid next_step answer");
  if (next.choice === "continue" && next.confidence >= 0.75) break;
  if (next.choice === "escalate") return handOffToReviewer(decision);
  await agent.gatherMoreEvidence(next.probabilities);
}

Validate the response

The response repeats your question names under answers. Each answer is a discriminated object whose type matches the requested question.

{
  "model": "weka",
  "answers": {
    "needs_review": {
      "type": "noul",
      "noul": 0.93
    },
    "release_lane": {
      "type": "choice",
      "choice": "canary",
      "confidence": 0.8,
      "probabilities": {
        "canary": 0.8,
        "stable": 0.16,
        "blocked": 0.04
      }
    },
    "risk": {
      "type": "score",
      "score": 1.54,
      "confidence": 0.72,
      "legend": {
        "0": "Low",
        "1": "Medium",
        "2": "High",
        "3": "Critical"
      },
      "probabilities": {
        "0": 0.08,
        "1": 0.42,
        "2": 0.38,
        "3": 0.12
      }
    }
  },
  "usage": {
    "input_tokens": 426,
    "output_tokens": 73
  }
}

Reject an unknown answer type, an option that was not offered, or missing question names. Log the input state version, answer, probabilities, and policy result when a decision changes production behavior.

Use Weka through the API, SDK, or Console

  1. Open the Console Playground.
  2. Select Weka decisions.
  3. Start from the release gate or incident triage example.
  4. Edit the JSON state and questions, then run the decision.
  5. Inspect the typed answer, confidence, probability distribution, and token usage.

The Autohand SDK and Console Playground use the same /v1/decisions contract as a direct API request. Every surface uses the selected account and consumes the same plan allowance.

JevBench v1.3.0 · 534 decisions per system

How Weka compares on JevBench

JevBench combines intelligence, calibration, speed, and cost at 25% each. Weka sits directly below Jev in this comparison of leading typed decision systems.

SystemScoreIntelligenceCalibrationSpeedCost
Jev 1.13.0 (TypeSafe AI)74.485.782.783.352.0
Weka (Autohand AI)74.689.580.281.349.0
SemIf (Qwen3.5-4B)73.179.072.683.759.5
djev (Maisa · diffusion-gemma)73.082.765.491.457.6
Winnow-12B Q871.282.072.082.352.9
reflex 4B70.380.175.268.059.7
jqv (Qwen3-32B zero-shot)68.679.379.074.647.5
decision-machine-168.362.170.492.953.7
decider-35b-a3b67.679.671.580.845.3
open-alternative-jev (Qwen3.5-4B)67.064.063.283.559.6
system-one-open (Gemma 4 E2B)66.669.556.777.064.8
OpenJev (DiffusionGemma 26B-A4B)66.479.264.883.245.5
SimpleJev (Qwen3.8-27B)66.384.781.171.239.5

JevBench v1.3.0 evaluates 534 decisions per system. The score is the geometric mean of the four axes. See the full leaderboard, public task results, and methodology.

Plans, quotas, and errors

Weka is included on Pro and above. It is not included on Free. Weka, Fantail, and Moa share the request allowance for the account; calling the public API does not bypass metering.

PlanWeka
FreeNot included; Free has a 200 request monthly hard stop
ProIncluded
MaxIncluded
TeamIncluded
EnterpriseIncluded

Autohand checks the account before calling Weka. A denied request returns HTTP 429 with the exhausted quota window, reset time, and retry metadata. Invalid state or questions return HTTP 400 before the request is metered.

Operational guidance

  • Keep AUTOHAND_API_KEY on a server or in a secret store.
  • Send the smallest state needed for the decision and omit secrets and personal data.
  • Use reviewed thresholds for production actions and keep a human path for low confidence results.
  • Version question instructions and criteria when decisions must be reproducible.
  • Use idempotency in the action layer so retries do not deploy or roll back twice.