Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
comgen42 
posted an update 3 days ago
Post
100
Kodiak v0.2 1B got its official Decision Index listing this week: 11.8, rank 96 of 115.

I'll be honest, that one stung. I've put a lot into this model. Our own run of the public benchmarks said about 19, and the index confirmed that part (19.1). But most of the full score comes from private tests, and on tasks from new domains Kodiak scored close to zero. It's good at what it was trained for: intent routing at 91.6, common sense, ranking. It falls apart on kinds of decisions it has never seen. Sarcasm on real tweets came out worse than random.

There's a particular kind of tired that comes from working hard on something and having a scoreboard tell you you're near the bottom.

Then I remember what I've learned from history. The Wright brothers came home from Kitty Hawk in 1901 convinced the published lift tables were wrong. Instead of quitting, they built their own wind tunnel and tested hundreds of wing shapes. James Dyson went through more than 5,000 prototypes before one worked. Failing wasn't the exception for the people who built things that mattered. It was the job. What separated them was that they kept going and kept measuring honestly.

So that's the plan. The index just told us exactly where the weakness is: generalizing to new kinds of tasks. Meanwhile a Hugging Face user found two shortcuts in v0.3, we fixed one and are testing the fix for the other right now, and every number goes into the public log, good or bad.

Rank 96 is a data point, not a verdict. Back to work.

github.com/grizzlypeaksoftware/kodiak

"Worse than random" on sarcasm is the wrong headline. I pulled Kodiak's row from the index's data/index.json.

Track A English is F1 on the sarcastic class, not accuracy. The 0.2227 random baseline implies about 1 tweet in 7 is sarcastic, so answering "yes" to every tweet would score about 0.25. Kodiak scores 0.189. A model with zero signal that says "yes" about 28% of the time lands right there. Below that baseline, F1 mostly measures how often you say yes.

Track C takes the threshold out: given a sarcastic tweet and its literal rephrasing, pick the sarcastic one. Kodiak is 0.505 on English pairs and 0.465 on Arabic. A coin is 0.5.

So the honest reading is "no sarcasm signal yet", not "reads sarcasm backwards". Those need different fixes.

Same shape elsewhere: 7 of your 42 benchmarks land below baseline and clip to 0 skill. RAGTruth's baseline is always answering "hallucinated", and 42 of 114 models lose to it. HLE has 108 of 114 below random, so that one is the benchmark, not Kodiak.

Your private-test drop (public 19.05 vs 11.8 overall, new domains 6.87) looks like the real story. Do you get per-item confidences back from the index, or only the aggregates?

·

You're right, and thanks. Track A is F1 on the sarcastic class; Kodiak says "yes" to 19.7% of the English tweets (276 of 1,400), so 0.189 is what no signal looks like there. Track C at 0.505 is the honest number: no sarcasm signal on real tweets yet. I've corrected our log.

On confidences: for the public suite, yes. We ran it ourselves and kept per-item calibrated probabilities for all 150,759 requests. For the private and new-domain tests we only see the totals, which is frustrating, since that's where the real drop is (new domains 6.87 raw).