Qwen3.5 27B vs Gemma 4 31B IT Here is why I think Gemma 4 31B wins: Benchmark scores of both models

Qwen3.5 27B vs Gemma 4 31B IT 

Here is why I think Gemma 4 31B wins:

Benchmark scores of both models are very close.
Generalization is not.

I reused NVIDIA’s CoDeC metric to probe "benchmark affinity".
Officially, CoDeC > 80 suggests contamination.
I've never seen this. But even below that, it can still reveal how benchmark-affine a model is.

Higher CoDeC scores mean the model is more "comfortable" with the benchmark.
So the ideal profile is:
high benchmark scores + low CoDeC.

By that lens, Gemma 4 looks clearly stronger than Qwen3.5.

Gemma 4 likely generalizes better to unseen tasks.
I observed the same pattern earlier with Gemma 3 vs Qwen3.

I wrote about this a few months ago:
https://kaitchup.substack.com/p/did-the-model-see-the-benchmark-during

Benjamin Marie的帖子配图 1:Qwen3.5 27B vs Gemma 4 31B IT Here is why I think Gemma 4 31B wins: Benchmark scores of both models
← 返回 XPut 首页 保存在 X 查看原帖 ↗