Is LLM the Steam Engine? A Calm Assessment for 2026
The analogy comparing LLM to a steam engine gained popularity in 2023. Three years later, the data shows it partially holds, but the probabilistic nature creates fundamental differences from the steam engine’s determinism. The core conclusion: LLM is a powerful cognitive tool whose value depends on how users harness its output.
Why This Analogy Matters
When James Watt improved the steam engine in 1769, no one foresaw it would spawn railways and factories. LLMs in 2026 may be at a similar early stage—but this “may be” is crucial. Data suggests that LLMs may be lowering the marginal cost of cognitive work, similar to how the steam engine lowered the marginal cost of physical labor. This hypothesis could be entirely wrong, and we will address this possibility specifically at the end.
What Was the Real Breakthrough of the Steam Engine
The core of the steam engine was not “doing more work” but three things: delocalization of energy, economies of scale, and lowered skill barriers. LLMs show similar potential characteristics: anyone can access tools like ChatGPT, one model can serve millions of users, and the barrier to entry for basic cognitive tasks may be lowered. But top cognitive talent remains highly concentrated, and the barrier may shift to “critical thinking” and “AI collaboration skills.”
What LLMs Can Do in 2026
We note five core capabilities that are currently being deployed.
Text Generation and Content Creation. Claude 4 supports 200K+ context; ByteDance’s Doubao has an estimated daily active user count exceeding 50 million. However, complex reasoning tasks still exhibit logical discontinuities, and factual accuracy cannot be guaranteed.
Code Assistance and Software Development. Cursor is valued at approximately $2.9 billion and has changed the workflow of millions of developers. GitHub Copilot is used by over 10 million developers globally, with a code acceptance rate of approximately 30-40%. But complex system architecture design still relies on humans.
Information Retrieval and Knowledge Q&A. Perplexity has an estimated monthly active user count exceeding 30 million, providing direct answers with cited sources. But citation accuracy remains a challenge, and most models cannot access the internet in real time.
Multimodal Understanding and Generation. GPT-4o and Gemini 2.5 Pro are natively multimodal, with improved image understanding accuracy. Video generation quality has improved significantly, but physical consistency remains a challenge.
Agent Autonomous Execution and Tool Calling. The MCP protocol has become infrastructure for the Agent ecosystem, supported by OpenAI, Google, and others. But autonomous planning capabilities are limited, and multi-step tasks frequently fail at intermediate steps.
Current Capability Boundaries: What LLMs Clearly Cannot Do
Data shows that the essence of LLM is “next token prediction,” not “fact-checking.” This means it may generate information that appears plausible but is completely incorrect, provide mutually conflicting answers within the same conversation, and cannot distinguish between fact and fiction.
Status as of June 2026: even the most advanced models exhibit hallucination rates varying from low single digits to over 20%, depending on the task, in fields requiring precise facts such as law, medicine, and finance. For high-risk decisions, the current error rate remains unacceptable.
Multi-step planning is another bottleneck. When a task requires a precise sequence of more than 5 steps, the failure rate rises significantly. Models cannot reliably check their own reasoning processes, and performance drops sharply when faced with entirely new problem types not present in the training data.
Actual Changes at the Societal Level
We observe several signals that are already unfolding.
Developer workflows have been transformed. Cursor and Claude Code have changed the daily coding practices of millions of developers, but the prediction that “AI will replace programmers” has not materialized. Coding efficiency may have improved, but architecture design, requirement understanding, and team collaboration still depend on humans.
Early signals are appearing in organizational structures. LinkedIn data shows that the number of “AI Product Manager” positions grew by over 200% year-over-year in 2025, but the absolute number remains small. Some startups demonstrate high leverage from small teams plus AI, but this is not yet a systemic trend.
The education sector faces pressure. The barrier to entry for programming has been lowered, but predictions of a “collapse of programming education” have not materialized. The actual change is that educational content has shifted from “syntax memorization” to “problem decomposition plus AI collaboration.”
Risks and Probabilistic Challenges
A steam engine explosion is a deterministic risk that can be controlled through engineering standards. LLM hallucination is a probabilistic risk that occurs randomly and is difficult to predict; traditional engineering methods cannot fully eliminate it.
In 2024, a New York lawyer used ChatGPT to generate a legal brief, citing fabricated case law, and was sanctioned by the court. From 2025 to 2026, the barrier to deepfake technology has dropped significantly, and AI-generated fake content in political elections has increased markedly. The failure rate of enterprise AI projects is estimated by Gartner at approximately 80% to 85%, which stands in stark contrast to success stories.
Three Possible Future Paths
Path One: Gradual Enhancement. LLMs continue to improve, but the Scaling Law may encounter bottlenecks. Capability boundaries expand slowly, and society adapts gradually. Assumption failure condition: if no major architectural breakthrough occurs by the end of 2027, or if inference cost reduction stalls.
Path Two: Agent Revolution. Autonomous Agent capabilities break through, transforming from assistants to executors. Assumptions are stringent: planning capabilities improve significantly, the tool-calling ecosystem matures, and reliability reaches production grade. Assumption failure condition: if Agent task success rates do not break through 80% during 2026-2027, or if hallucination rates cannot be reduced below 1% for critical tasks.
Path Three: Winter and Consolidation. Technical bottlenecks plus regulatory tightening plus investment cooling push the industry into a consolidation phase. Assumption failure condition: if a major architectural breakthrough occurs, or if AI application ROI is validated at scale.
Why This Analogy Might Be Wrong
We note five counter-narratives.
Cognitive Fireworks, Not Cognitive Steam Engine. The Transformer architecture may be approaching its performance ceiling, high-quality text data may be nearing depletion, and OpenAI and Anthropic are estimated to face massive losses. The killer app has yet to emerge, and most predictions from 2023-2024 have not materialized.
Cognitive Opium, Not Cognitive Steam Engine. Excessive reliance on AI may lead to the degradation of human critical thinking; users may gradually lose the ability to fact-check. The standardization of AI-generated content may inhibit the diversity of human creativity.
Cognitive Arbitrage, Not Cognitive Steam Engine. The low cost of LLMs comes from compute subsidies and venture capital; if priced at true cost, many applications may not be economically viable. The premium for high-quality human cognitive work has not declined due to AI; on the contrary, it has risen due to the flood of AI content.
Wrong Time Scale. From Watt’s improvement of the steam engine to the completion of industrialization took over 100 years; from the invention of electricity to the completion of electrification took over 40 years. It has only been 6 years since the release of GPT-3, and society already expects a revolution. This compression of the time scale may be unrealistic.
Cognitive Catalyst, Not Cognitive Energy. LLMs do not produce cognition; they merely accelerate the flow of existing cognition. In domains where AI improves efficiency, costs in other domains often rise.
Actionable Advice for You
Based on the capability boundaries as of June 2026, at the individual level we recommend choosing a professional domain to deepen your expertise while mastering AI tool usage. At the organizational level, we recommend shifting from replacing human labor to augmenting human labor, embedding AI into workflows while preserving human review checkpoints. At the societal level, a balance between innovation and safety is needed, but premature strict regulation may stifle innovation.
An Open Question
If LLM is neither a steam engine, nor fireworks, nor opium—if it is an entirely new technological form—then what framework should we use to understand it? Historical analogies may be a useful thinking tool, but not a reliable predictive framework. The most likely scenario is that LLM requires its own analytical framework, rather than historical analogy.