Model News

Nace AI Open-Sources Drex 1.5 Decision Model for Option Scoring

Nace AI open-sources Drex 1.5, a 9B decision model that scores multiple options in one forward pass, ranking first under 10B parameters on Decision Index.

Headline card: Nace AI Open-Sources Drex 1.5 Decision Model for Option Scoring
On this page
  1. What changed
  2. Why it matters
  3. What to test
  4. The conclusion

What changed

Nace AI has open-sourced Drex 1.5, a 9B decision model built for agent pipelines and backend workflows. The model reads text or JSON state alongside typed questions and returns a probability for every option in a single forward pass. Unlike general-purpose language models, Drex 1.5 does not generate text, code, or explanations.

The model serves the POST /v1/systemone API, the same format used by TypeSafe’s Jev model. Nace reports that existing Jev clients can switch to Drex 1.5 by changing environment variables. Weights are available on Hugging Face under a RAIL-M license, with a hosted version live on OpenRouter at $0.04 per 1M input tokens and $0 for output tokens.

Drex 1.5 uses MiMo-V2.6-Distill-Qwen-9B as its backbone, a distilled Qwen 3.5 9B model with 32 layers and hybrid attention. The model adds a separate pointer head that scores each option from the backbone’s hidden states. The context window expands from a default of 16,384 tokens to 131,072 tokens, a significant extension for processing long documents.

Why it matters

Decision models serve a specific function in production systems: they classify, route, or score predefined options without the overhead of text generation. Traditionally, teams use classification APIs or hosted services for these tasks. Drex 1.5 changes the economics by offering open weights that run locally on a single GPU or even on consumer hardware via quantization.

The benchmarks matter here. Nace reports Drex 1.5 scores 58.08 on Decision Index 0.3.1, a public 37-benchmark suite. This ties with Jev 1.13.0 at 57.96, within the leaderboard’s 0.9-point tie band. Both beat Bespoke Nimble 9B v3 at 57.19 and Cloudflare clef-flash at 56.15. Drex leads on 20 of 37 benchmarks.

For long documents, the case is stronger. On the public JevBench, Drex scores 86.2% against Jev’s 87.0%. But where Drex stands out is on extended contexts: 93.4% accuracy on 32K to 128K token documents versus 89.5% on 8K to 32K tokens. Truncating the same requests to 8K tokens drops accuracy to 76.5% and 78%. This suggests Drex 1.5 is optimized for contracts, policy reviews, and document classification workflows.

For practitioners, the implications are practical. Teams running agent systems with routing decisions, tool selection, or escalation logic can now avoid API calls and keep sensitive data on their own infrastructure. The model fits on a single A10G GPU in bf16 precision and shrinks to 9.5 GB as a Q8_0 quantized GGUF that runs on Macs and CPUs. Nace offers deployment paths through Python Kev runtime, llama.cpp forks, Ollama forks, and OpenRouter hosting.

What to test

Before adopting Drex 1.5, verify these vendor claims and limitations:

Benchmark scope

Nace reports that Drex was trained on the official Decision Index training splits and evaluated only on held-out splits. Test the model on your own domain-specific decision tasks. Knowledge-heavy tests show weakness: Drex scores 45.4% on GPQA Diamond versus Jev’s 78.6%, and 58.7% on MMLU-Pro versus 82.7%. If your options involve reasoning over unfamiliar domains, results may not match index numbers.

Long-document accuracy

Nace claims 93.4% accuracy on 32K to 128K token documents. Reproduce this on your actual document types and lengths. The company notes median latency of 2.0 seconds at 32K to 128K tokens, but test on your hardware and quantization choice.

Compatibility and forks

Drex runs through Nace forks of llama.cpp and Ollama, not mainline builds. Confirm these forks work with your deployment pipeline. The RAIL-M license has use restrictions; review the terms for commercial applications.

Sentiment and fine-grained classification

The model shows weak performance on ACOS aspect sentiment at 7.4% per-review F1, versus Jev’s 29.5%. If your use case involves nuanced sentiment or multi-aspect classification, test against your own labels first.

API compatibility

Nace claims the model works with the /v1/systemone API. Test this API in your environment to confirm request and response formats match your orchestration code.

The conclusion

Drex 1.5 delivers on a practical promise: a small decision model that performs competitively with its closed counterpart on a public leaderboard, runs locally, and handles long contexts well. The 93.4% accuracy on 32K to 128K token documents is a genuine strength for document triage and contract review. The tie with Jev on Decision Index 0.3.1 suggests parity for typical agent routing tasks.

The weak spots are real. Knowledge-heavy tasks expose the model’s limits, and sentiment classification shows significant drops versus Jev. These are not surprises given Drex is purpose-built for decision scoring, not general reasoning. The dependency on Nace forks for llama.cpp and Ollama adds friction for teams using mainline tools.

The open weights matter most. Teams handling sensitive data in regulated domains can now keep decision logic on their own infrastructure without sending documents to hosted APIs. The core value is infrastructure control and cost predictability for high-volume pipelines.

What to watch next: whether Nace updates Decision Index performance on version 0.4.x benchmarks, how the model performs on sentiment and knowledge tasks in real-world deployments, and whether mainstream llama.cpp and Ollama projects upstream Nace’s decision model features. The model’s strength on long contexts also positions it for document classification workflows beyond agent routing, making it worth evaluating for any workflow that requires scoring predefined options over extended text.

AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure