Repository navigation
SystemOne Decision Models Integration into lemonade #3751
meghsat
started this conversation in
Request for Comment (RFC)
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Proposal size
Major feature (spans multiple well-scoped PRs)
Updates
Implemented this here: #3752
User story
As a developer I want to get decisive, straightforward, and quick responses to my questions and LLMs[system-two] models are slow and often produce way too many tokens unless we prompt tune them. To solve this, a new set of models called system-one is released. These models are built to make fast, structured decisions that software can use directly.
llama.cpp added decision models in PR #29818. The PR was merged on 2026-10-02, and release b11361 contains it. A decision model answers typed questions in one forward pass. It does not generate tokens. llama-server serves these models through the /v1/systemone endpoint. The GGUF files are released for the following models:
These models can also be extended into the router colelction.router policy. As a developer, I want to ask a small local model typed questions about each request. Then I can route a request on rules that I write in plain words. For example: "Does the message contain personal data?" or "Which team must handle this?".
High-level design
Lemonade gives decision models to clients and to the router in three ways.
POST /v1/systemoneThe endpoint is also available under
/api/v0/,/api/v1/,/v1/, and/v0/. Lemonade does these steps:Request:
Response (Laya, values rounded):
POST /v1/classifyA decision model also answers a
/v1/classifyrequest that has labels. Lemonade asks the model one choice question, with the labels as the options. The response has the same form as the response of the ONNX zero-shot model. Thus the zero_shot router classifier from PR #3740 operates with a GGUF decision model, and the router does not change.The recipe of the model selects the backend:
flowchart LR C1[Client] -->|POST /v1/systemone| H1[handle_systemone] P[collection.router policy] -->|systemone| RS[Router::systemone] H1 --> RS P -->|zero_shot| RC[Router::classify] C2[Client] -->|POST /v1/classify + labels| RC RC -->|onnxruntime| O[ort-server] RC -->|llamacpp| A[LlamaCppServer::classify<br/>one choice question] RS --> L[LlamaCppServer::systemone] A --> LS[llama-server /v1/systemone] L --> LSA
systemoneclassifier asks one typed question in your words. Its labels are the options of the question:choicecriterianoultrueandfalsescorecriteria{ "classifiers": [ {"id": "pii", "type": "systemone", "model": "Laya-GGUF", "default_label": "true", "question": {"type": "noul", "instructions": "Does the message contain personal data?"}}, {"id": "topic", "type": "zero_shot", "model": "Laya-GGUF", "labels": ["chit-chat", "coding"]} ], "rules": [ {"id": "pii-local", "match": {"classifier": "pii", "min_score": 0.5}, "route_to": "Local-GGUF"}, {"id": "code-to-big", "match": {"classifier": "topic", "label": "coding", "min_score": 0.5}, "route_to": "Big-GGUF"} ] }Lemonade adds these models to the registry:
Julia-1-GGUFlaya(encoder)Q8_0Laya-GGUFlaya(encoder)Q8_0Kev-4B-GGUFkev(causal)Q8_0lev-GGUFlev(causal)Q8_0OpenJev-GGUFopenjev(causal, images)Q4_K_M+mmprojThe
classificationlabel is the deployment mode. Thus a decision model never loads as a chat model, and it does not unload the chat model of the user. Thesystemonelabel tells Lemonade that the model is a decision model. A llama.cpp model with theclassificationlabel and without thesystemonelabel is still refused at registration.Implementation
/v1/systemoneinllama-serverzero_shotclassifier type,labelson/v1/classify/v1/systemone, the/v1/classifyadapter, thesystemoneclassifier type, the launch settings, the two models, tests, and the specdocs/dev/specs/systemone-decision-models.mdBreaking changes
None. The endpoint, the classifier type, and the two models are additions.
Maintenance plan
CI tests the feature. Upstream llama.cpp maintains the prompt layout and the scores of each decision model.
/v1/systemone,/v1/classifyadapter,systemoneandzero_shotclassifierssystemone-llamacpp(cpu, Julia-1 and Laya) and C++ unit testsRisks
No response
All reactions