Instructions to use q-project/Q-U-164M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use q-project/Q-U-164M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="q-project/Q-U-164M", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("q-project/Q-U-164M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use q-project/Q-U-164M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "q-project/Q-U-164M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "q-project/Q-U-164M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/q-project/Q-U-164M
- SGLang
How to use q-project/Q-U-164M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "q-project/Q-U-164M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "q-project/Q-U-164M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "q-project/Q-U-164M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "q-project/Q-U-164M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use q-project/Q-U-164M with Docker Model Runner:
docker model run hf.co/q-project/Q-U-164M
This is a research release, not a product. It is small and often factually wrong outside its trained tool-calling domain. See Results below before using it for anything beyond experimentation.
Model summary
Q-U-164M is Q-164M โ a 164.6M-parameter ternary-weight,
frozen-fingerprint-vocabulary text model โ fine-tuned in two stages specifically for dialogue: chat-SFT,
then RL on a verifiable reward supplied by the model's own deterministic tool executor (the same
<CALC>expr<EQ> mechanism the base model uses, circuits.py). Same architecture and weights format as the
base model, different weights.
Current release: v5. The first four fine-tuning passes all shared the same real bug: in a multi-turn conversation the model would drop tool-calling by turn 2-3 and drift into disconnected prose. The actual cause only surfaced with a controlled A/B test โ the synthetic training conversations never included conversational connectives (โthanks, now what's...โ, โone more question...โ), just bare questions back to back, so the model had never seen โacknowledge small talk and still call the toolโ in one training example, and real users write exactly that way. v5 splices connectives directly into the training data. Full story: blog post.
Results
| capability | result |
|---|---|
| single-turn tool-calling probe (6 fixed tasks) | 6/6 |
| 8-turn multi-turn stress test, natural phrasing | 6/8, zero repetition-collapses across the entire training run |
| RL training stability | KL never exceeded 0.004 across 200 steps โ the most stable fine-tuning run yet |
Known remaining weaknesses, shown honestly: occasionally confuses adjacent operations (e.g. division with multiplication) with a correct-looking but wrong tool expression; rare misclassification of open-ended, non-numeric questions into the wrong tool.
How it compares to other small models
Measured with QBench โ no model gets forced-correct tool execution in this comparison, including this one, so scores are comparable across architectures. Live table: huggingface.co/spaces/q-project/QBench.
| model | params | single-turn tool-use | multi-turn tool-use | ARC-Easy | TruthfulQA MC1 |
|---|---|---|---|---|---|
| GPT-2 | 124M | 0% | 0% | 42% | 26% |
| SmolLM2-360M-Instruct | 360M | 50% | 50% | 50% | 18% |
| Qwen2.5-0.5B-Instruct | 500M | 45% | 50% | 62% | 24% |
| Q-U-164M | 164M | 35% | 25% | 43% | 26% |
Q-U-164M doesn't win outright here, and shouldn't be expected to โ it's a third the size of Qwen2.5-0.5B, trained on one V100. The numbers above give no model any tool-execution help, for a fair comparison. Deployed normally โ with its own deterministic circuit executor correcting whichever calculation it decides to make (as in the Results table above) โ its actual tool-call accuracy is much higher; the architecture's whole premise is that the network never has to get the arithmetic right itself, only decide when to ask.
Usage
The control tokens (<user>, <model>, <eot>, ...) are registered as named special tokens, so the
standard chat template works:
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("q-project/Q-U-164M")
model = AutoModelForCausalLM.from_pretrained("q-project/Q-U-164M", trust_remote_code=True)
ids = tok.apply_chat_template(
[{"role": "user", "content": "What is 91 divided by 7?"}],
add_generation_prompt=True, return_tensors="pt",
)
out = model.generate(ids, max_new_tokens=60)
print(tok.decode(out[0], skip_special_tokens=True))
For guaranteed-exact tool calls (forces the deterministic result instead of trusting the network's own
arithmetic), use circuits.CircuitLogitsProcessor โ see the
GitHub repo for a full runnable example
(examples/chat.py) and the fine-tuning recipe that produced this checkpoint (finetune/).
Architecture
Same as the base model โ 10 layers, hidden 1280, 20 query / 4 KV heads, ternary weights, frozen 512-bit fingerprint vocabulary. See the base model card for full architecture details.
Limitations
- Facts are frequently wrong outside its trained tool-calling domain; thin general knowledge. No safety tuning.
- Occasionally confuses adjacent tool operations (see Results).
- Ternary weights at this size are a research setting; do not expect competitive benchmark scores from a single-V100, 164M-parameter model.
License and data terms
Code and weights: Apache-2.0 (see LICENSE). Fine-tuning data keeps its own terms โ check before
commercial use: GSM8K (MIT), SmolTalk's everyday-conversations subset (Apache-2.0); the rest of the
fine-tuning mix is synthetically generated by this project.
- Downloads last month
- 88
Model tree for q-project/Q-U-164M
Base model
q-project/Q-164M