Small models.
Real decisions.
Give a model a question. Inspect every score. Then put it in the arena.
Your decision appears here.
— server msInspect response + receipt
A small model we trained.
Editable categories scored by a byte-level attention model, trained locally on synthetic support-routing examples.
192-byte input; 32 bytes per option. New labels are accepted, but competence beyond the six training categories is not established.Learned option scores
— ms including tokenization and headThis is a newly trained checkpoint—not a general replacement for GLiClass or Jev. Try wording it has never seen and inspect the result.
A model in the control loop.
Your own isolated arena. Two arenas available globally. Up to 60 game seconds / 180 wall seconds.
Waiting for observation
0KILLS
—HEALTH
0GAME SECONDS
GLiClass classifies text derived from visible engine objects. Code chooses targets and controls movement. Not pixel vision or learned pathfinding. Simulation pauses for inference.
Measure it. Don’t take our word for it.
—
Accuracy on 120 synthetic test cases with unseen templates but shared vocabulary.
8 / 10
Stock NLI backend on ten authored adversarial cases. It failed missing-evidence and revocation examples.
Small development test—not an independent benchmark.Three identified backends.
GLiClass, MiniLM NLI, and our trained option-attention checkpoint. A unified interface—not one general reasoning model.
No verified Jev comparison. No synthetic latency claims.Training report
Raw inference scores are not calibrated probabilities of correctness. The tiny model’s held-out templates do not establish generalization to real customer records.