NextFin News — For most of the past three years the loudest competition in large language models centered on scale: more parameters, more data, more cloud compute. A quieter parallel track focused on the opposite problem—how to deliver useful capability inside the memory, power and thermal limits of consumer devices. That track is now intersecting with real product cycles.
![]()
Smartphones, cars and other endpoints increasingly require AI that can function with limited or no cloud connection. Local execution reduces latency, protects certain categories of data and shifts inference cost onto hardware the customer has already purchased. The result is growing demand for models that are small enough to fit, efficient enough to run continuously, and robust enough to survive multi-year product lifetimes.
One illustration is ModelBest, the Beijing company behind the MiniCPM series. Founded in 2022 from a Tsinghua University laboratory, it concentrated early on efficiency rather than maximum size. Researchers associated with the firm later described an empirical pattern they called a density law: under their measurements, the capability extracted per parameter roughly doubled every few months. By mid-2026 the open-source MiniCPM line had recorded tens of millions of downloads. The company’s cabin software has entered production vehicles from several Chinese carmakers, with management projecting hundreds of thousands of cars equipped by year-end. Reports have also linked the models to upcoming flagship phones.
The commercial structure of on-device AI differs sharply from cloud APIs. Revenue is not a clean multiple of tokens consumed. Models must be adapted to specific system-on-chip platforms, validated through long automotive or consumer-electronics development schedules, and supported across years of software updates. Once a solution is integrated and certified, the cost for an original-equipment manufacturer to switch suppliers tends to be higher than the cost of changing a cloud endpoint. That stickiness is a potential advantage for the model provider. It is also a constraint: the OEM controls the final product, the user relationship and most of the pricing power. How much of the value created by local intelligence flows to the model specialist is a matter of negotiation, not automatic scaling.
Customer concentration adds another layer of risk. A handful of large vehicle or phone programs can deliver rapid volume. The same concentration can leave a supplier exposed if any one relationship changes or if the next generation of hardware favors a different stack.
Competition has broadened quickly. In 2024 specialized on-device models were still relatively uncommon. By 2026 major device makers and cloud providers were all shipping or preparing compact models optimized for local execution. Size alone is no longer scarce. Any lasting differentiation is more likely to rest in training efficiency, runtime software, chip-specific optimization, and the ability to support volume production and long-term maintenance.
Regulators have begun to reflect the industry’s stage of development. Updated guidance on China’s STAR Market allows certain artificial-intelligence model companies to seek listings even if conventional revenue thresholds have not yet been met, provided they can demonstrate at least one model in scaled commercial use. The rule change acknowledges that technical maturity and commercial scale are not always synchronized.
The deeper question for the sector is how the economic value of on-device intelligence will be divided. Chip vendors, operating-system owners, device makers and model specialists all contribute. End customers ultimately pay for the finished product. Early volume proves that compact models can reach mass hardware. It does not automatically determine which layer of the stack captures durable margin.
Private markets have already assigned high valuations to companies that sit at the model layer of this stack. Public markets, once they open, will demand clearer evidence of pricing power, recurring software revenue after initial integration, and resilience to customer concentration. The shift from technical demonstration to supply-chain reality is under way. Whether the specialists who arrived first can retain a meaningful share of the resulting value remains the open test for the industry.










