No. 1
Can a Local Model Read My Inbox? A Bake-off
I gave three on-device models the same job - turn each email into a structured record - and judged them against Claude Opus. Three models, four runs, because Qwen ran once with thinking on and once with it off. With thinking off, the smallest-feeling model won on faithfulness, classification and speed.
- task
- Turn each email into a structured record: a summary, a type, an importance from 1 to 5, and any actions or dates
- judge
- Claude Opus 4.8 in the cloud, as the quality ceiling
- contestants
- Qwen3.6-35b-a3b (thinking off)
- Qwen3.6-35b-a3b (thinking on)
- Gemma-4-12b-it (8-bit MLX)
- Gemma-4-e4b-it
- winner
- Qwen3.6-35b-a3b (thinking off)
- axes
- faithfulness, classification, importance, speed
- verdict
- Local on-device inference is good enough for this job - the winner sweeps speed and quality. The judge was a cloud model, so the test emails were read in the cloud. In daily use a model on my own machine reads the mail, and no cloud model does.