Q-U-164M logo

Q-U-164M

Q-164M, fine-tuned for dialogue with exact tool calls

Parameters Ternary weights License

Chat demo Blog GitHub

This is a research release, not a product. It is small and often factually wrong outside its trained tool-calling domain. See Results below before using it for anything beyond experimentation.

Model summary

Q-U-164M is Q-164M โ€” a 164.6M-parameter ternary-weight, frozen-fingerprint-vocabulary text model โ€” fine-tuned in two stages specifically for dialogue: chat-SFT, then RL on a verifiable reward supplied by the model's own deterministic tool executor (the same <CALC>expr<EQ> mechanism the base model uses, circuits.py). Same architecture and weights format as the base model, different weights.

Current release: v5. The first four fine-tuning passes all shared the same real bug: in a multi-turn conversation the model would drop tool-calling by turn 2-3 and drift into disconnected prose. The actual cause only surfaced with a controlled A/B test โ€” the synthetic training conversations never included conversational connectives (โ€œthanks, now what's...โ€, โ€œone more question...โ€), just bare questions back to back, so the model had never seen โ€œacknowledge small talk and still call the toolโ€ in one training example, and real users write exactly that way. v5 splices connectives directly into the training data. Full story: blog post.

Results

capability result
single-turn tool-calling probe (6 fixed tasks) 6/6
8-turn multi-turn stress test, natural phrasing 6/8, zero repetition-collapses across the entire training run
RL training stability KL never exceeded 0.004 across 200 steps โ€” the most stable fine-tuning run yet

Known remaining weaknesses, shown honestly: occasionally confuses adjacent operations (e.g. division with multiplication) with a correct-looking but wrong tool expression; rare misclassification of open-ended, non-numeric questions into the wrong tool.

How it compares to other small models

Measured with QBench โ€” no model gets forced-correct tool execution in this comparison, including this one, so scores are comparable across architectures. Live table: huggingface.co/spaces/q-project/QBench.

model params single-turn tool-use multi-turn tool-use ARC-Easy TruthfulQA MC1
GPT-2 124M 0% 0% 42% 26%
SmolLM2-360M-Instruct 360M 50% 50% 50% 18%
Qwen2.5-0.5B-Instruct 500M 45% 50% 62% 24%
Q-U-164M 164M 35% 25% 43% 26%

Q-U-164M doesn't win outright here, and shouldn't be expected to โ€” it's a third the size of Qwen2.5-0.5B, trained on one V100. The numbers above give no model any tool-execution help, for a fair comparison. Deployed normally โ€” with its own deterministic circuit executor correcting whichever calculation it decides to make (as in the Results table above) โ€” its actual tool-call accuracy is much higher; the architecture's whole premise is that the network never has to get the arithmetic right itself, only decide when to ask.

Usage

The control tokens (<user>, <model>, <eot>, ...) are registered as named special tokens, so the standard chat template works:

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("q-project/Q-U-164M")
model = AutoModelForCausalLM.from_pretrained("q-project/Q-U-164M", trust_remote_code=True)

ids = tok.apply_chat_template(
    [{"role": "user", "content": "What is 91 divided by 7?"}],
    add_generation_prompt=True, return_tensors="pt",
)
out = model.generate(ids, max_new_tokens=60)
print(tok.decode(out[0], skip_special_tokens=True))

For guaranteed-exact tool calls (forces the deterministic result instead of trusting the network's own arithmetic), use circuits.CircuitLogitsProcessor โ€” see the GitHub repo for a full runnable example (examples/chat.py) and the fine-tuning recipe that produced this checkpoint (finetune/).

Architecture

Same as the base model โ€” 10 layers, hidden 1280, 20 query / 4 KV heads, ternary weights, frozen 512-bit fingerprint vocabulary. See the base model card for full architecture details.

Limitations

  • Facts are frequently wrong outside its trained tool-calling domain; thin general knowledge. No safety tuning.
  • Occasionally confuses adjacent tool operations (see Results).
  • Ternary weights at this size are a research setting; do not expect competitive benchmark scores from a single-V100, 164M-parameter model.

License and data terms

Code and weights: Apache-2.0 (see LICENSE). Fine-tuning data keeps its own terms โ€” check before commercial use: GSM8K (MIT), SmolTalk's everyday-conversations subset (Apache-2.0); the rest of the fine-tuning mix is synthetically generated by this project.

Downloads last month
88
Safetensors
Model size
0.2B params
Tensor type
F32
ยท
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for q-project/Q-U-164M

Base model

q-project/Q-164M
Finetuned
(1)
this model

Collection including q-project/Q-U-164M