Qwen3.5 27B vs Gemma 4 31B IT Here is why I think Gemma 4 31B wins: Benchmark scores of both models
原帖发布于
Qwen3.5 27B vs Gemma 4 31B IT
Here is why I think Gemma 4 31B wins:
Benchmark scores of both models are very close.
Generalization is not.
I reused NVIDIA’s CoDeC metric to probe "benchmark affinity".
Officially, CoDeC > 80 suggests contamination.
I've never seen this. But even below that, it can still reveal how benchmark-affine a model is.
Higher CoDeC scores mean the model is more "comfortable" with the benchmark.
So the ideal profile is:
high benchmark scores + low CoDeC.
By that lens, Gemma 4 looks clearly stronger than Qwen3.5.
Gemma 4 likely generalizes better to unseen tasks.
I observed the same pattern earlier with Gemma 3 vs Qwen3.
I wrote about this a few months ago:
https://kaitchup.substack.com/p/did-the-model-see-the-benchmark-during

Here is why I think Gemma 4 31B wins:
Benchmark scores of both models are very close.
Generalization is not.
I reused NVIDIA’s CoDeC metric to probe "benchmark affinity".
Officially, CoDeC > 80 suggests contamination.
I've never seen this. But even below that, it can still reveal how benchmark-affine a model is.
Higher CoDeC scores mean the model is more "comfortable" with the benchmark.
So the ideal profile is:
high benchmark scores + low CoDeC.
By that lens, Gemma 4 looks clearly stronger than Qwen3.5.
Gemma 4 likely generalizes better to unseen tasks.
I observed the same pattern earlier with Gemma 3 vs Qwen3.
I wrote about this a few months ago:
https://kaitchup.substack.com/p/did-the-model-see-the-benchmark-during
