Overfit in the field · iPhone · unlisted

Gemma 4 on iPhone

26B parameters. One phone. No cloud.

A 26B-parameter mixture-of-experts model whose full weights do not fit in an iPhone’s memory, answering a multi-step ballistic-pendulum problem on the device itself. Overfit keeps the shared layers resident and pages routed experts from local storage as the router asks for them — so the model runs at its real size, at storage speed rather than memory speed.

Prefill speed
34.4tk/s
Prefill time
20.34s
Decode speed
3.5tk/s
Noema on iPhone showing a multi-step ballistic pendulum physics problem sent to gemma-4-26B-A4B-it-Q4_K_XL, with the model status bar reading READY and 1.9k of 4.1k context used.
01
The prompt

A ballistic-pendulum problem chained across a spring compression and a pendulum swing. Model loaded and resident at 1.9k / 4.1k context.

The same conversation scrolled to the end, showing the momentum-conservation working and a final answer of 661 metres per second, with run statistics of 1,201 tokens at 3.5 tokens per second, 20.34 seconds to first token and 362.45 seconds total.
02
The answer

1,201 tokens of working, ending at 661 m/s. The run footer reports the on-device timings, and the badges confirm the whole thing stayed local.

Exactly what was on the device.

Model
gemma-4-26B-A4B-it — Q4_K_XL
Parameters
26B total · ~4B active per token
Device
iPhone — no Mac, no relay, no cloud
Execution
Overfit paged experts from local storage
Context
1.9k of 4.1k used
Output
1,201 tokens · 362.45s wall clock

One capture on one device, not a benchmark. Prefill and decode speed depend on storage throughput, available memory, quantization, prompt length, expert reuse, and thermal state — a smaller model that fits entirely in memory will still start and generate faster.

Models larger than memory. Still local.

How Overfit works