← home

AI competition.

Pushing the boundaries of a small language model.

2nd Solo8th Team

competition

Improve the math reasoning capabilities of a 4B model using post-training and inference-time techniques — no external APIs, tools, or second model at inference time.

TEST SET
943

Maths questions

REASONING MODEL
Qwen

Qwen3–4B

SCORE
74.5%

Accuracy

COMPUTE / RECORD

The burn — by the numbers

<think>213,000,000</think>

reasoning tokens generated

generated
5,309responses
GPU inference
258hours

THE DATASET

943 math questions.

MCQone answermulti answer
300305338

sources

SuperGPQA · UGMathBench

licence

CC BY-NC-SA 4.0

THE MODEL Qwen3–4B.
mode

Thinking.

<think>
8 can't be right. I'll check the ratio. Let me try graphing it.
</think>