The authors used Intelligence per Watt as a metric to evaluate the efficiency of local AI inference, and found that local models under 20b active parameters can successfully answer 88.7% of single turn chat and reasoning queries. Between 2023 and 2025, the intelligence efficiency of these models improved by a factor of 5.3 due to advances in both model architectures and hardware accelerators.

Local models are now capable of handling the vast majority of everyday user requests without relying on any centralized cloud infrastructure. While cloud models are still better at highly specialized reasoning, deploying small models on personal devices is quickly becoming a practical and energy efficient alternative.