All postsTag
Benchmarks
2026
Lab
Can a Local Model Read My Inbox? A Bake-off
I gave three on-device models the same job - turn each email into a structured record - and judged them against Claude Opus. Three models, four runs, because Qwen ran once with thinking on and once with it off. With thinking off, the smallest-feeling model won on faithfulness, classification and speed.
A general-purpose MoE multimodal beat every dedicated vision model on my father's handwriting
I assumed a specialized vision model would win. I was wrong. A head-to-head on a hard handwriting corpus ended with the general-purpose MoE on top.