Oscar-1: Decisions, Not Text
Give it a state and typed questions; it returns typed answers with calibrated probabilities and a confidence score. No text generation, so nothing to parse and nothing to hallucinate.
| type | question | answer |
|---|---|---|
| choice | which of these options? | the option, a probability for each, confidence |
| score | where on this rubric? | a position along your levels, probabilities, confidence |
| noul | is this true? | the probability that it is |
Every tab answers all of its questions in one forward pass, and the action line is plain code reading those
numbers: the thresholds live in the app, not in the model. Pick a checkpoint below — 17m
and 32m are RLCD-trained Ettin deciders, trained with proper
scoring rules (log + spherical + RPS), laya.Agent-compatible, 2.5–2.6 ms p50 on GPU.
Preview checkpoints: 17M and 32M parameters, trained with RLCD on a 59,135-item variants+mix corpus. Trained strengths: banking intents (0.81–0.86), news topic (0.87), support triage (0.66–0.70), moderation-style flags (0.73). Weak, treat as untrained: guardrail screens (0.29–0.46), movie-review sentiment (0.28–0.46), NLI (0.35–0.39), any multilingual input (0.23–0.28) — the tabs still run so you can watch the calibration behave, but route real traffic to stronger checkpoints.
Running on CPU. ZeroGPU could not attach a GPU for this Space (
Error), so answers take a few hundred milliseconds instead of ~35 ms. Everything still works.
Classify, detect urgency, score frustration and check for a refund request in one call, then route.
answers
12 banking intents — the trained strength of these checkpoints (0.81–0.86 on the sealed harness).
answers
News topic + sentiment rubric + dominant emotion in one pass — topic 0.87, the mix-corpus strengths.
answers
Quoted replies, signatures and disclaimers are stripped in code first, then one call answers five questions. These checkpoints saw no email data — this tab is generalisation, not a trained skill.
answers
Screen prompts before the expensive model: jailbreak, injection, sensitive data, harm. Weak on these checkpoints (0.29–0.46) — watch the calibration behave, don't trust the verdicts.
answers
Score retrieved passages for relevance, contradiction and hidden instructions; keep what earns its place. Passage relevance is not a trained skill on these checkpoints — expect flat, low-calibrated scores.
ranked passages
answers
Grade difficulty, domain and tool need, then send each request to the cheapest model that can handle it.
answers
Checkpoints: mgoeckel/oscar-1-17m, mgoeckel/oscar-1-32m · ZeroGPU · weights loaded and warmed at start-up
Layout and demo patterns adapted from convaiinnovations/laya-demo (Apache-2.0).