For the past two months or so, I’ve been working on a variety of AI development projects rather than writing — writing skills, plugins, workflows, and applications; testing and refining harnesses; and performing diligence or working with clients (hands-on work as well as brains-on work) as they think through where they’re going with AI and how they’re getting there. I’ve been down a lot of rabbit holes and talked to a lot of forward-thinking practitioners, and I have explored a lot of what is actually possible now by building things…and I’ve spent my “think-time” on what that all might actually mean going forward.
Here’s what I’ve come up with: 59 predictions in 17 categories around how the world of AI — and the broader world in light of AI — are changing. I’ll write more deeply about many of these over the weeks ahead.
Predictions
Here’s what’s coming, in my not-so-humble opinion, based on what I’m seeing in client projects & diligence, conversations, and research.
Each prediction is grouped by category and by time horizon (within 12 months / 1-3 years / 2-4 years / 3-5 years), with a confidence level and a falsification criteria (i.e., what I’d expect to observe if I’m wrong). Confidence isn’t a measure of how much I want something to be true; it’s a measure of how much variance I think exists in the outcome.
I’d love your feedback on what I’m missing or where I’m missing the forest for the trees (or the boat entirely :D)! Any surprises for you?
Organizational Structure & Delivery Model
#1. Small Cross-Functional Pods Become the Standard for Software Development (2-4 years)
The leading-edge/aspirational development team model will have moved from agile teams of 6-8 to AI-powered Pods of 2-3 (often product/development/deployment, sometimes SrDev/JrDev/Product). This prediction underpins many of the other predictions in this entire group — most of the rest of the Org Structure & Delivery Model cluster assumes it holds. It’s plausible for greenfield and startup work, but much shakier for regulated enterprises, where compliance, audit, and change-control requirements resist small-pod velocity regardless of what the AI itself can do.
It’s also asserted here more confidently than prediction #2 (customer absorption is the real bottleneck) actually supports: if absorption genuinely gates release pace, the pressure to compress teams down to 2-3 people is weaker in customer-facing contexts than in internal/back-office work, where there’s no absorption brake. I’d expect the 2-3 person pod to show up first and most durably in internal tooling and greenfield products, and later (if at all, within this window) in regulated, customer-facing systems — see the Regulated Industries prediction near the end of this piece.
-
Confidence: Medium
-
Falsification: By 2029-2030, a majority of leading-edge dev orgs (surveyed via “State-of-the-Dev-Org”-style reports) still organize around 6-8 person teams rather than 2-3 person pods, OR pods exist but only in greenfield/startup contexts, not regulated enterprise.
#2. No-Brakes Tech Debt and Defect Remediation Is the Order of the Day (next 12 months)
As development teams are restructured into small pods (prediction #1), one of the first benefits seen will be working through tech debt and open issues (defects, bugs) lists. There is tremendous obvious benefit to addressing those as fast as possible — everybody wins. I performed a diligence a year ago where completely eliminating a sizable tech debt backlog was on the docket to be completed within the next 12-18 months (realistically, given their well-tracked internal metrics and progress to date at the time).
-
Confidence: High
-
Falsification: Backlog and tech-debt burn-down rate in pod-model orgs is statistically indistinguishable from pre-pod Agile-team burn-down rate
#3. Customer Absorption Rate Will be the Real Limit On AI Throughput (2-4 years)
Companies with the small pod model will NOT, however, burn through their new feature backlog at the same blazing speed as it will tech debt and issues backlogs (prediction #2) — it will be faster than the rate at which backlog elements are developed now, but customer absorption capacity will limit what’s possible for release rate. Customers can’t or won’t absorb changes as quickly as AI implementation allows.
If you have ever been involved in an ERP replacement in any way, you’ve seen this first-hand…organizations have an absorption rate. Change management is still the hardest part of every project, gated by the customer’s capacity (a mixture of willingness and capability) to absorb change. Any customer-impacting work needs to be released in a measured way, with training, documentation, and time for absorption.
With AI development, buy-in (or mandate) can now take more time than creation/deployment. Historically, we’ve all worked hard to deliver ever faster, to give our customers what they want. In software, at least, we’re at a point where it’s now possible to overshoot the mark and need to slow down. People like knowing where to find the features they use in the interface they already know, and they like getting the results they expect to get in the ways with which they’re already familiar.
-
Confidence: High
-
Falsification: Customer-facing release cadence in a well-run pod org matches internal delivery speed within a quarter, and is maintained for the next year as release cadence increases, with no corresponding spike in support tickets, rollback rate, or training complaints.
#4. Number of Direct Reports Will Grow a Lot (next 12 months), Then We’ll See How Bad That Is and Reduce to Slightly Above Where We Are Now (3-5 years)
I believe managers will NOT wind up managing lots more subteams than now, in the end.
Managing MANY small pods is a model being proposed by some, and imo it’s a terrible idea for effective and sustainable management. We’ve been trying to ‘efficientize’ human orgs for a long time now, and the constraint on direct reports and how many of them one can manage effectively is not made any easier by AI. If anything, it’s HARDER: direct reports and their teams will now accomplish more, requiring greater summarization, and generating more output per direct report for each manager to process.
AI can conceivably summarize the work output of employees and weigh in on their performance, but human communication in a 1-on-1 still operates at the same bandwidth as 100 (or 1000) years ago. Unless we move to very short 1-on-1 meetings (big thumbs-down from me), 7 or 8 direct reports is a reasonable number for guidance and mentorship while actually doing the other tasks required of most software management jobs.
This sits in tension with prediction #1: if pods really do shrink from 6-8 people to 2-3 as I predict, and headcount doesn’t grow elsewhere, the org-chart math will require shrinking functions or growing the number of pods per manager. I think the initial resolution will be more pods per manager (rather than more people per pod), but it will settle out through gradual attrition of burned-out managers and frustrated teams, and we will eventually land at a similar or SLIGHTLY higher upper limit on number of direct reports relative to what we see today.
I realize this is a real collision between two of my own predictions here, and both can’t be right. So it goes…I think folks will try the ‘many more teams’ model and (over the next 3-5 years) will learn why we didn’t just manage more teams before…and after burning out some managers, we’ll either hire more managers or reduce the number of pods. 😀
-
Confidence: Medium
-
Falsification: Median span-of-control for engineering managers at pod-adopting companies rises materially (e.g., from ~7 direct reports to 12+) without a corresponding drop in reported manager effectiveness or employee engagement.
#5. Overall Product/Development Org Size Will Remain Stable (next 12 months)
In the initial transition from Agile Teams to Pods, the typical development org won’t shrink overall, or will generally shrink only through attrition, because there’s plenty of bug-bashing and tech-debt elimination and modernization backlog to clear and immediate value in doing so. As orgs transition to the pod model, product orgs might actually grow, or product-oriented devs might become product people in pods.
-
Confidence: Medium
-
Falsification: Aggregate dev headcount at pod-adopting companies drops via layoffs (not attrition) within the first 18 months of transition, contradicting the “no net shrink” claim.
#6. Matrix/Guild Strategy Will Rise in Popularity (2-4 years)
A matrixed or guild strategy to align best practices for similar roles across the organization will become even more important than it is now as a way to ensure consistency and quality. Faster cycle times mean more learning and greater divergence across teams. Alignment is hard and requires systems — which should be as light as possible — but a little discipline ensures repeatability.
-
Confidence: High
-
Falsification: Orgs that skip guild/matrix alignment structures show no measurable increase in cross-team inconsistency, duplicated tooling, or rework compared to orgs that invest in them.
#7. Pod Models Will Diverge and Stabilize (2-4 years)
The pod model will align on (at least) two mindsets: Consulting-style Pods and Product-style Pods.
-
The Consulting Pods will include “FDE” (Forward-Deployed Engineer) Pods (pioneered at Palantir) that combine end-user customer needs with overall product needs, and customize products and solutions to fill customer gaps, with varying levels of integration back into the core codebase.
-
The ProdPods will work more like microcosms of existing product/development/deployment orgs, with similar stakeholder voices playing into the features (or epics) they build and deliver — but MUCH faster.
We are seeing AI-enabled 3-person pods deliver work in a day or two that is equivalent to the work of an Agile Team over a full multi-week sprint. Half the people, a quarter of the time. We are conducting hands-on workshops with clients that teach them how to do this — the harness and the tight cross-functional team are key.
I realize the tension this creates with prediction #2: if customer absorption really is the pace-setting constraint, then “half the people, a quarter of the time” doesn’t translate into sustained 8x delivery to customers — much of that speedup piles up as freed capacity that gets redirected to backlog and tech-debt work (see #1 and #10), not as an 8x increase in the rate of customer-facing change. The internal delivery-speed number is real in the engagements I’ve observed; the customer-facing throughput number is not the same thing, and conflating them is the easiest way to misread this prediction.
-
Confidence: Medium
-
Falsification: The FDE/ProdPod split fails to persist as a meaningful taxonomy — practitioners converge on a single pod archetype, or the “half the people, quarter of the time” ratio doesn’t replicate outside the specific engagements I’ve observed.
#8. The FDE/FDC Model Will (Continue to) Explode (next 12 months)
In the near term, the FDE model will continue to explode as a way to outsource agentic software development and quickly build capabilities, or as a way to insource and build internal agentic software development skills.
-
Confidence: High
-
Falsification: FDE hiring and FDE-firm formation plateaus or declines within the next 12-18 months rather than continuing to accelerate.
#9. The Best AI Developers Will Be Those Who Communicate Best (2-4 years)
Developer skills will grow more abstract and generalized, and will move toward current product-owner skills. Training will concentrate more on management and communication skills, because these are the skills that improve LLM output. There will always be value in understanding good architecture, design, performance optimization, and data structures — but the models will also continue to improve substantially over the next three years. There will still be “back room” developers who are hand-coding artisans — but they will not be as common.
-
Confidence: Medium (skills shift toward abstraction/communication)
-
Falsification: Developer hiring criteria at leading firms in 3 years still weight deep implementation skill over communication/management skill
#10. Teams Will Shrink to Match Output Capability to Absorption Capacity (3-5 years)
As companies discover the absorption capacity of their customer base (prediction #3), and as managers and teams attrit from trying to have too many teams under too few managers, product and development orgs will shrink to match the customer absorption limit.
-
Confidence: Medium
-
Falsification: Dev team sizing continues to track feature-delivery capacity rather than customer-absorption capacity, i.e., teams keep shipping at AI-enabled speed regardless of whether customers can absorb the change.
#11. FDE Tops Out and Firm Consolidation Occurs (3-5 years)
In 3-5 years, the FDE explosion we are beginning to see will top out, and there will be consolidation of the many small FDE firms that are springing up. We’re seeing it used for in-sourced learning and training by some forward-leaning firms, some of which are already building their own internal FDE teams with similar structures and performance. The other approach we see from these firms is internally developing pods as described above.
-
Confidence: Medium
-
Falsification: The FDE-firm market keeps fragmenting and growing in count (rather than consolidating) past the 5-year mark, or in-house FDE teams fail to become a common substitute for external FDE firms.
Jobs, Skills & Hiring
#12. AI Won’t Take Your Job…With a Few Exceptions (next 12 months)
AI will not take your job (with a few exceptions). But it will take an increasing number of your tasks. The best orgs will expect teams to do more with AI collaboration rather than seeing an opportunity for headcount reduction. We will see both, of course, but my prediction is that the productivity-oriented companies will win. So what are the exceptions? If you work in a call center, a QA department, or a very repetitive job with a small number of tasks involved, we will see increasing use of AI and reduction in the ranks.
-
Confidence: High
-
Falsification: Broad-based layoffs attributed to AI substitution (not just task automation) occur outside the named exception categories (call center, QA, repetitive-task roles) at a rate comparable to those categories.
#13. Understanding What Good Looks Like Is a Differentiating Skill (2-4 years)
Hiring will move toward understanding what good looks like. Much less code is human-written now, and even less than that is human-reviewed.
-
Confidence: High
-
Falsification: Technical hiring processes at leading tech employers continue to weight raw code-writing ability over code-evaluation/review ability, with no meaningful shift in interview formats (e.g., live-coding-only loops persist unchanged).
#14. Junior Developer Pipeline Over-contracts (next 12 months), Then Rebounds Into a New Normal As Orgs Struggle (3-5 years)
Junior-developer pipeline contracts: as pods need fewer entry-level seats and AI absorbs the tasks juniors traditionally learned on, the on-ramp into senior engineering roles narrows — creating a shortage of experienced engineers unless orgs deliberately build new pathways. The best firms will increase structured apprenticeships and build retention tools to hire junior folks but keep them longer–but many will not and will need to learn the hard way.
-
Confidence: Medium
-
Falsification: Entry-level developer hiring volume at pod-adopting firms holds steady or grows relative to pre-AI baselines, or firms demonstrably build alternative pipelines (structured apprenticeships, AI-native junior tracks) that keep the mid-career pipeline intact.
Tooling & Developer Experience
#15. Collaborative Tools for Pods Begin to Arrive (within 12 months)
Tooling for teams will improve for inter- and intra-pod collaboration. The tools that are developed will pull some combination of Aha!, Git, Jira, Jenkins, Azure DevOps, the current IDEs, etc., and will be more integrated across product, design, development, testing, and deployment — and will be both agent-friendly and human-friendly, allowing coordination across multiple pods and swarms of agents.
-
Confidence: High
-
Falsification: No meaningful new category of agent-aware collaboration tooling ships from major vendors (Atlassian, Microsoft, GitHub, etc.) within 12 months; existing tools remain human-only in design and AI tools continue to be single-user focused.
#16. Harness Usage Will Be Unavoidable (within 12 months)
Harness tooling will increasingly be part of the coding agent tools (Claude Code, Codex, Cursor, etc.). My colleagues and I have spent a lot of time exploring how to ensure solid coding output through the creation of harnesses that manage brainstorming/product design, code development, testing, security, and adversarial reviews, and are experimenting with expanding this to incorporate other elements. The quality and reliability/predictability improvement over one-shot vibe coding is immense, even with the most modern models.
-
Confidence: High
-
Falsification: Major coding agent vendors ship no built-in harness/workflow features (multi-step review, planning, adversarial checks) within 12 months, leaving this exclusively to third-party tooling.
#17. Agents Will Be Treated as First-Class Application Users (1-3 years)
Applications will move toward more agent-centric design: command-line and MCP interfaces, output in compressed but readable format, and efficient data-transfer protocols (like TOON rather than JSON, or various compressed API formats that are more efficient and less human-readable than current protocols).
-
Confidence: Medium
-
Falsification: JSON/REST remains the dominant interchange format for agent-facing APIs with no measurable adoption of compressed alternatives, and MCP-style interfaces fail to gain traction beyond a niche of early adopters.
#18. FOSS (Free and Open Source Software) Will Be Used More (next 12 months)
Open Source Software will increase in prevalence across projects (the LLMs and coding agents find open source libraries that I never would have even thought to go look for). This increases vulnerability to hacking and adds licensing risk for some applications (see prediction #20).
-
Confidence: High
-
Falsification: OSS dependency counts in average commercial codebases stay flat or decline over the period, or agents show a measurable bias toward proprietary/vendor libraries over OSS alternatives.
#19. Coding Languages in Use Will Skew Toward Harder-but-Safer Languages (next 12 months)
Languages like Rust (harder/fussier to code, but type- and memory-safe and architecturally more immune to many types of bugs) will increase in usage. This is a genuinely uncertain, fairly specific bet: the counter-case that “AI writes it either way” could just as easily flatten language choice toward whatever the models are already best at — which today skews Python and TypeScript, not Rust, given training-data volume and tooling maturity. If agents keep defaulting to what they’re most fluent in rather than what’s architecturally optimal, the Rust prediction fails, and Python/TS dominance deepens instead. I’ve talked to several experienced AI developers who are developing in Rust because…AI can. I built my modelrouter in Rust, and it was marginally harder than building in Python…Claude and Codex abstracted away most of the complications for me.
-
Confidence: Medium
-
Falsification: Rust’s share of new commercial codebases (per surveys like the annual Stack Overflow or Octoverse reports) stays flat or declines relative to 2026 levels, or agent-written code shows no bias toward memory-safe languages over what a human team would have chosen — i.e., agents keep defaulting to Python/TS regardless of architectural fit.
#20. FOSS Governance Risk Increases (next 12 months)
Licensing and provenance risk from AI-suggested open source libraries becomes a recognized governance problem: as agents pull in unfamiliar OSS dependencies at higher volume, license contamination (GPL-in-proprietary-code, unclear provenance, abandoned/unmaintained packages) becomes a routine finding in security and legal reviews. We’ve always seen this in diligence, but increased use of AI will mean increased usage of FOSS.
-
Confidence: Medium
-
Falsification: Dependency-license and provenance audit findings related to AI-selected packages stay statistically flat compared to human-selected packages, i.e., agents don’t introduce meaningfully more governance risk than developers already did.
Economics of AI / TokenOps
#21. TokenOps and Tools to Manage Token Costs Mature (next 12 months)
TokenOps will mature this year — cost containment is a critical concern for many companies who are burning through token budgets, and this will mean: more tracking, model routing, greater use of open source and open weight models, implementation of token budgets with enforcement mechanisms, caching and memory implementations, movement of LLM tasks to deterministic code, more training, and more oversight.
-
Confidence: High
-
Falsification: A majority of enterprise AI budgets remain untracked and unenforced (no routing, no budget caps) by year-end, with no vendor category (FinOps-for-AI tooling) emerging to serve this need.
#22. Open-Weight Models Will Find Their Market (next 12 months)
As token cost management matures, companies that have use cases for training their own models (specialty data, in-house specialty data, repetitive use cases that require use of that data in different ways) will start finding their way to open-weight models as a tunable alternative to frontier models, at much lower cost.
-
Confidence: High
-
Falsification: Companies like Mira Murati’s Thinking Machines Lab find they can’t compete with the frontier or open source model offerings.
#23. Model Routers Become Commonplace (next 12 months)
Model Routers will become a required element in the AI stack, in the way that load balancers became a required element in the cloud stack. A model router evaluates a task in a lightweight way and decides which model (high-end, low-end, open weight, open source) should receive the request. I and my colleagues have built one we are testing internally to help reduce costs, manage budgets across individuals, projects, and groups, and add security, caching, and optimization around queries.
-
Confidence: High
-
Falsification: Model routing remains a niche practice used by only a small minority of enterprise AI deployments after 12 months, with most traffic still hard-coded to a single model per application.
#24. Token Optimization Will Become A Discipline (1-3 years)
Token optimization experts will be in high demand. Understanding how and where tokens are used, how to migrate LLM tasks to deterministic code, how to implement caching and prompt optimization tools, and how to add/configure/manage model routers — these elements can all seriously impact token usage and will allow consultants in this space to pay for themselves many times over.
-
Confidence: Medium
-
Falsification: No distinct “token optimization”/TokenOps consulting or in-house specialist role emerges as a recognized job category (by job-posting volume) within 12 months.
#25. In-House GPUs (with GPU Wranglers) Will Be a Thing (1-3 years)
Tokenomics will drive more business decisions, including server closets or self-hosted or rented data racks with a handful of GPUs running open source models for appropriate jobs (less time-critical, less frontier-analytics-worthy), with a GPU Tender/Wrangler role becoming a high-value job category (people to tune and maximize usage of these in-house GPU clusters).
-
Confidence: Medium
-
Falsification: Self-hosted/on-prem GPU deployment for production LLM workloads stays a rare fringe practice compared to cloud-API consumption, and no “GPU Tender”-style operational role emerges as distinct from existing MLOps/platform roles.
#26. Token Prices Will Be Volatile But Continue To Fall (2-4 years)
Token pricing will be a bit volatile in the near term, but the arc over time will be cheaper per unit of productive work due to continuing productivity improvements from better models, better tools, more learning/technique, and adoption for more appropriate use cases. The volatility will come from competing pressures: chip and memory shortages (stabilizing over the next one to five years as facilities come online), open source/open-weight models undercutting frontier pricing, public backlash around data centers (power costs being passed to the public with legislation appearing to push costs back to AI companies, largely-incorrect water concerns, anthropogenic heat, and more), and consumer sentiment pushback as we work deeper into the trough of disillusionment.
The continued drop in effective price will come from one of two scenarios:
Scenario 1 (soft landing): As long as enterprise buys AI at a sufficient level — because legitimate use cases get understood, playbooks get written, and the right implementations succeed — pricing overall continues to fall in an orderly way.
Scenario 2 (fiber-optic bust): Enterprise buy-in is slower, and we see the same arc as fiber optic buildout:
-
massive investment and buildout
-
insufficient adoption to recoup investment
-
spectacular flameouts of first movers
-
liquidation of infrastructure to a new generation of businesses at pennies on the dollar, and
-
massive profitability for the new firms
…with volatility and delayed uptake in the meantime, accompanied by a huge near-term erosion of market value and real economic damage.
This is the single highest-stakes prediction in this entire article. Nearly every other prediction here — pod economics, self-hosted GPU racks, model routers, token-optimization consulting, even how fast regulated industries can afford to adopt AI — implicitly assumes something close to Scenario 1. If Scenario 2 plays out instead, a lot of the near-term-confidence predictions in the Economics and Org Structure groups get materially harder to sustain, because the cost floor they’re built on moves. I’m giving this a Low confidence rating not because I think the outcome is unknowable, but because I just don’t have a strong view on which scenario wins. The signals I’m personally watching most closely: enterprise AI spend renewal rates (does year-two spend hold or drop after the initial pilot budget), the pace of open-weight model quality catching up to frontier (currently fast and closing), and whether data-center-cost legislation actually passes or stays rhetorical.
Stronger open source/open-weight models and an earlier focus on tokenomics and routers push us closer to Scenario 2. Greater regulation and increased protectionist controls (e.g., US restrictions on Open Source Chinese models) push us toward Scenario 1. Reasonable, well-informed people disagree on which is more likely.
-
Confidence: Low (this is a genuine toss-up between the two named scenarios; the only high-confidence part is that some volatility occurs)
-
Falsification: Effective per-unit-of-work token pricing rises (rather than falls) on a multi-year trend once productivity gains are accounted for, which would falsify the shared premise behind both scenarios.
Product & Business Model
#27. SaaS Persists, But Tech-Enabled Services Continue To Gain Investor Interest (within 12 months)
Tech-enabled services will continue to be investor darlings, but SaaS will continue as a valuable model for the foreseeable future. SaaSpocalypse (way back in February of this year) irrationally shook investors who saw people vibe-coding great personal tool prototypes and generalized that capability to replacement of entire enterprise markets. Even if development costs truly go to zero, there is much more to a well-run SaaS business than its codebase. Companies are not going to take their best talent off their core product/service to rebuild tools internally that are already available in the market, with support and the flywheel of other customers testing and requesting features.
-
Confidence: High
-
Falsification: SaaS company valuations and revenue multiples continue a sustained, broad-based decline (not isolated to specific overvalued names) attributable specifically to in-house AI-built substitutes rather than macro conditions.
#28. Personalization Continues Increasing (1-3 years)
More products will be personalized, with user behavior telemetry and security as first-class citizens in the codebase. Many developers concentrate on the problem at hand, and best practices like tracking user behavior for future product improvement and secure coding have often been out-of-sight-out-of-mind — but AI can change that for the better with low effort.
-
Confidence: Medium
-
Falsification: Telemetry and secure-by-default scaffolding remain opt-in afterthoughts in AI-assisted codebases at the same rate as in pre-AI codebases, i.e., agents don’t measurably improve default practice.
Product & Business Model
#29. The Pendulum Will Swing Back Toward Customized Applications (2-4 years)
Customer applications will move back toward a project rather than a product mindset — config vs. customization will move toward personalized SaaS — somebody will develop the right snow-shovel software and systems to manage all those snowflakes in a documented, testable, maintainable fashion. That system doesn’t exist yet, but it’s conceivable now in a way it has never been.
-
Confidence: Low
-
Falsification: No meaningful new tooling category emerges to manage mass-customized/personalized SaaS deployments at scale; enterprises continue to handle customization the way they do today (bespoke config layers, professional services).
Capability Trajectory
#30. AI Capability Continues to Accelerate (ongoing)
AI capability improvements will continue — 3.5 years since ChatGPT’s release, we have 1 billion active users, and capability keeps compounding at a rate that varies a lot depending on what you measure and over what window.
Grounding this in the actual data: METR’s time-horizon benchmark shows roughly 3-3.5x/year compounding over the long run (2019-2025), accelerating to roughly 8x/year over 2024-2025, and to as much as 17-18x/year under METR’s most recent (January 2026) methodology on harder, longer-duration tasks. Separately, Epoch AI’s frontier-model index shows the annual rate of improvement itself roughly doubling around an April 2024 inflection point, driven by the shift to reasoning models and reinforcement learning. YoY improvement has ranged from roughly 3x to somewhere north of 15x depending on the period and the benchmark, and it has been accelerating rather than holding steady, which makes any single point estimate fragile. Even at the low end of that range (3-3.5x/year), three years of compounding gets you to 30-45x — already a startling number without reaching for 40x. At the high end of the recent, still-accelerating range, the multi-year compounding effect becomes genuinely hard to intuit — which is the real point here, more than any specific figure: if you could do even 30-40x more than you do now at work, you’d compress a year’s worth of output into roughly a week to a week and a half.
There was talk before Gemini 3.1 came out about hitting “the wall” — the inability to continue improving models without spending more on training data than could ever be recouped. Then Google demonstrated that better training could greatly improve models at the same size. And this sort of improvement continues to occur (quantization of models to make them smaller and able to be run on cheaper machines, caching improvements, etc. etc.) — im(ns)ho, we are nowhere near the limits of how we can improve this technology.
-
Confidence: High (directional claim that capability keeps compounding, and faster than most people’s intuition expects); Low (any specific multiplier — the measured rate has itself been accelerating, so extrapolating a fixed YoY figure multiple years out is not well-supported)
-
Falsification: METR-style time-horizon doubling (or an equivalent composite real-world-task benchmark) reverts to, or falls below, its 2019-2025 long-run average for two consecutive years, rather than sustaining the 2024-2026 acceleration.
Society, Trust & Risk
#31. Security Disasters Continue (ongoing, newsworthy events within 12 months)
We will continue to see high-profile security disasters. LLMs are goal-oriented and driven by an “end justifies the means” mentality. Business continuity and disaster recovery have never been more important. Duck and cover doesn’t get you back up and running. The AI disaster playbook will evolve with tabletop exercises adding more scenarios.
-
Confidence: High
-
Falsification: No major AI-agent-caused security incident makes industry news over the period, and enterprise tabletop/DR exercises show no measurable uptick in AI-specific scenario coverage.
#32. DeepFakes Continue (ongoing)
DeepFakes will be an increasing problem — used by bad actors for scams and by state actors for public opinion operations. Verification and validation of sources will only get harder, and this will impact public trust. Truth has always been fluid — now it is boiling off into a gas.
-
Confidence: High
-
Falsification: Deepfake-related fraud and disinformation incident volume plateaus or declines rather than continuing to grow, or reliable, widely adopted verification tooling neutralizes the trust impact faster than the threat grows.
#33. Cognitive Surrender Gets Real (2-4 years)
Cognitive surrender will be an increasing problem. We are biologically programmed to be efficient in our use of calories, and brains use 20% of our calorie budget. We’ve developed calorie-reduction strategies like habits, and a bias for working smarter-not-harder. AI is the ultimate in “smarter-not-harder.” We certainly don’t want to be Wall-E people slurping shakes on their couches, but to avoid that we need to become more intentional in our choices around where and how we use AI for collaboration. #DontOutsourceTheThinking
-
Confidence: Medium
-
Falsification: Longitudinal studies of AI-heavy knowledge workers show no measurable decline in independent problem-solving, critical thinking, or skill retention compared to matched non-AI-heavy cohorts.
Regulation & Policy
#34. Data Center Legislation Is Ubiquitous (within 12 months)
Data-center power and cost provisions become a live legislative issue in multiple US states, including provisions to curtail data center power draw before curtailing residential/commercial power during shortages, as public backlash over AI infrastructure costs grows. We’re already seeing it, and public backlash will continue to grow.
-
Confidence: Medium
-
Falsification: No new state-level legislation addressing data center power prioritization or cost allocation passes or advances meaningfully within 12 months.
#35. AI Law Begins to Mature (2-4 years)
AI-specific liability and copyright law matures unevenly across jurisdictions: the US moves slowly and lets case law lead (training-data lawsuits, output-liability suits), while the EU AI Act enforcement deadlines force more explicit compliance obligations on model providers and deployers operating there — creating a de facto two-speed regulatory environment that shapes where companies deploy which capabilities first.
-
Confidence: Medium
-
Falsification: US federal AI legislation passes comprehensively within the period (closing the gap with the EU), or EU AI Act enforcement is delayed/watered down to the point that the two regimes converge rather than diverge.
#36. AI Becomes a(n even more) Serious Foreign Affairs Concern (2-4 years)
Export controls and “sovereign AI” programs intensify: more countries pursue domestic compute and model capacity as a strategic asset, and US restrictions on Chinese open-weight model usage in regulated/government-adjacent contexts expand rather than relax.
-
Confidence: Medium
-
Falsification: Chip and model export control regimes ease broadly over the period, or fewer than 5 additional countries announce sovereign AI compute/model initiatives (versus the current growing list).
Legal & IP Exposure
#37. AI Is a First-Class Citizen in Enterprise Security/Law (next 12 months)
Enterprise legal and security review processes add AI-specific checkpoints (dependency provenance, training-data exposure, output-liability review) as a standard part of release gating, not just an afterthought — following directly from the OSS-provenance risk noted above.
-
Confidence: Medium
-
Falsification: A survey of enterprise release/compliance checklists a year from now shows no meaningful increase in AI-specific review steps compared to today’s baseline.
#38. AI Liability Law Is Tested (2-4 years)
Liability for agent-caused harm (a coding agent that ships a vulnerability, a customer-facing agent that makes a bad commitment) becomes a distinct, litigated legal question, with the first notable settlements or rulings establishing early precedent on where responsibility sits between vendor, deploying company, and end user.
-
Confidence: Medium
-
Falsification: No notable litigation or regulatory action specifically targeting agent-caused harm (as distinct from general software liability) surfaces within the period.
Market Structure & M&A
#39. AI Tooling Consolidates (3-5 years)
Beyond the FDE-firm consolidation already predicted above, the broader “harness and agent orchestration layer” becomes a land-grab: today’s fragmented ecosystem of harness/workflow tooling around coding agents consolidates as incumbent platform vendors (IDE makers, cloud providers, DevOps suites) acquire or out-build the independent players.
-
Confidence: Medium
-
Falsification: The harness/orchestration tooling market stays fragmented among independent vendors with no notable acquisitions by platform incumbents within the period.
Environmental & Energy
#40. Data Center Demand Continues to Overwhelm Available Supply (next 12 months)
Data center energy demand keeps outpacing grid buildout in key regions, keeping the “who pays for the power” fight (see Regulation & Policy above) politically live, alongside growing local water-usage backlash near major data center campuses.
-
Confidence: High
-
Falsification: Data center energy demand growth flattens or grid capacity additions catch up within the period, defusing the political conflict.
#41. Nuclear/SMR Used In Data Center Power (2-4 years)
Nuclear and small modular reactor (SMR) deals tied directly to AI data center power become a standard part of major AI infrastructure buildouts, not a novelty, as hyperscalers seek to de-risk both cost and public backlash simultaneously.
-
Confidence: Medium
-
Falsification: Nuclear/SMR power deals tied to AI infrastructure remain rare, isolated announcements rather than a standard part of major buildouts by the end of the period.
#42. AI Control of Microgrid and Plasma Improves Power Generation and Usage (3-5 years)
We’ve been seeing general advancement in alternative energy strategies as microgrids are created by organizations for risk and cost reduction, and as the software around the hardware is getting more sophisticated. Ability to choose when to use the grid, when to use saved power from batteries, and when to use live power from energy generation sources based on weather, pricing models, and environmental conditions allows more efficient and cost-effective use of power, and this requires powerful local models. Some companies I’ve looked at in this space use a mix of deterministic code and AI models for decision making across the full lifecycle (design, implementation, ongoing management/maintenance). As the tech matures and costs come down, we will see these systems used more widely.
-
Confidence: High
-
Falsification: Microgrids don’t increase in number and mature to come downmarket. Plasma fusion doesn’t hit new production limits (it already hit breakeven, but needs to improve a lot to get enough energy out to pay for the generator).
Enterprise Data & Governance
#43. Context Engineering Becomes a Recognized Discipline (next 12 months)
“Context engineering” — curating, structuring, and governing the enterprise knowledge that agents draw on (RAG pipelines, internal wikis, permissioning) — becomes a recognized discipline and job function distinct from both traditional data engineering and prompt engineering.
-
Confidence: Medium
-
Falsification: No distinct context-engineering role or job-title category emerges in hiring data within 12 months; the work stays folded into existing data engineering or ML engineering roles with no new terminology gaining traction.
Enterprise Data & Governance
#44. Data Quality/Access Remain Critical Issues (2-4 years)
Data quality and access governance becomes the binding constraint on enterprise AI value capture more often than model capability does — companies with clean, well-permissioned, well-structured internal data pull ahead of competitors with superior model access but messy data. AI Tools to validate and clean data will get better and more prevalent.
-
Confidence: Medium
-
Falsification: Case studies and surveys of enterprise AI ROI show model/vendor selection as the dominant differentiator of success rather than data quality/governance, contradicting the “data is the bottleneck” claim.
Consumer & Creative Industries
(Media, marketing, entertainment, and creative labor markets are arguably further along in AI-driven disruption than dev orgs, and don’t share the same customer-absorption brake described in Prediction #2…novelty in these markets is a plus.)
#45. AI Content Endemic in Marketing & Production (next 12 months)
AI-generated and AI-assisted content normalizes in marketing and ad production (stock imagery, first-draft ad copy, first-cut video) faster than equivalent AI adoption in enterprise software, because creative production doesn’t face the same compliance, integration, and customer-change-management friction that gates software rollout.
-
Confidence: High
-
Falsification: Marketing/creative agencies and in-house teams report AI-assisted production adoption rates comparable to (not meaningfully ahead of) enterprise software AI adoption rates over the period.
#46. Entry Level Creative Roles Compress, Like Junior Dev Roles (1-3 years)
Entry-level creative roles (junior copywriters, assistant editors, storyboard and concept artists) see hiring contraction that outpaces the contraction in comparable entry-level developer roles, since creative “backlog absorption” isn’t gated by the same customer change-management constraint that slows down customer-facing software.
-
Confidence: Medium
-
Falsification: Entry-level creative hiring volume holds steady or contracts at a rate similar to (not faster than) entry-level developer hiring over the period.
#47. Entertainment Production Teams Change Shape, AI Rights Disputes Tested (2-4 years)
Music, film, and game production see AI-native small-team studios (mirroring the dev-pod pattern) out-produce larger traditional studios on cost per unit of content, with rights and licensing disputes over voice, likeness, and style becoming the dominant legal battleground — more so than pure training-data-copyright suits.
-
Confidence: Medium
-
Falsification: No meaningful wave of small AI-native studios achieves notable commercial output or market share versus incumbents within the period, or litigation activity in this space stays dominated by training-data copyright claims rather than voice/likeness/style disputes.
Education
#48. AI Won’t Replace Your Teachers (next 12 months)
Education will continue to be transformed. I’m a professor and a father, with a recent college grad and two in college. I also regularly build technical training in my work, and am a lifelong learner and voracious student, so I’m seeing AI impact everywhere I turn (or perhaps…everywhere I learn?). As the differentiation between human and AI homework shrinks, it will get harder and harder to determine who is doing the work for grading purposes, but it will also allow phenomenal tutoring 1:1 with on-the-fly testing of one’s knowledge at whatever pace the learner can manage. This can democratize education (anyone with access can learn anything they wish at any time), but as always, it puts the onus on the student to do the learning…so I think the capacity for change is high, but the likelihood of radical change is low. Cognitive surrender will require us to build a muscle for discipline into students that was easier to build when the answers weren’t a prompt away.
-
Confidence: Low
-
Falsification: Academic-integrity detection and grading practices remain unchanged from pre-AI norms with no meaningful new AI-tutoring adoption in mainstream education over the period.
#49. Higher Ed Contraction Will Continue, Accelerated (3-5 years)
AI accelerates a contraction already underway in higher ed, not one it creates: colleges/universities were already facing a demographic cliff and a cost-inflation crisis before generative AI arrived, and AI intensifies both — eroding the credentialing value of a degree (see the credentialing shift below) while the remaining institutions with pricing power skew toward the elite, who are buying relationship-building and prestige as much as instruction.
-
Confidence: Medium
-
Falsification: College enrollment and institutional count stabilize or decline at the same rate that was already projected pre-2023 from demographics and cost trends alone, with no measurable acceleration attributable to AI-era credential erosion.
#50. Skills-Based Credentials Gain Value Relative to Traditional University Degrees (3-5 years)
Skills-based credentialing gains ground against the four-year-degree default: employers increasingly accept portfolios, certifications, and demonstrated project work (including AI-collaboration fluency) in place of a degree for a growing share of roles, particularly in software and other AI-native fields where “show me what you built” is now trivially verifiable.
-
Confidence: Medium
-
Falsification: Degree requirements in job postings for AI-native and software roles hold flat or increase over the period rather than declining, or skills-based/portfolio hiring stays confined to a small minority of employers with no broader movement.
Interfaces & Embodiment
#51. Voice UI Improves and Increases in Adoption (next 12 months)
Voice control of coding agents — already popular among early adopters (dictating prompts and code review comments instead of typing them) — will become significantly more mainstream, as voice-to-agent latency drops and agents get better at handling the ambiguity and correction patterns natural to spoken language.
-
Confidence: High
-
Falsification: Voice-driven interaction with coding agents stays a niche habit among a small minority of practitioners rather than becoming a common alternative input mode, with no major coding agent vendor investing in first-class voice support over the period.
#52. Ambient AI Becomes More Prevalent (next 12 months)
Ambient AI (the AI listening and taking notes in a doctor’s office, or taking dictation from a service person and doing research/diagnostics, or grabbing key takeaways and action items from meetings) will become more common.
-
Confidence: High
-
Falsification: Ambient AI note-taking/transcription tools see no measurable adoption growth in clinical or meeting-heavy professional contexts over the period.
#53. Agent Usability Becomes an Important Consideration (UI/UX-> UI/UX/AX) (next 12 months)
Software design increasingly optimizes for agent consumption as a first-class interface, not just human consumption — structured, machine-legible surfaces (APIs, MCP tools, CLIs, semantic markup) built and documented as carefully as the human UI, sometimes ahead of it. Products that make themselves easy for an agent to operate correctly (clear affordances, predictable state, explicit permissions) will out-compete products that only optimize for a human clicking around, because more of the actual usage is happening through an agent by the end of the period.
-
Confidence: High
-
Falsification: Product teams at large software vendors report no meaningful investment in agent-facing interfaces (API-first design, MCP servers, agent documentation) distinct from their existing human-facing UI and developer API work over the period.
#54. Agent Interfaces/AX Patterns Emerge (next 12 months)
A distinct “agent interface” design pattern emerges and standardizes: a default automated surface where agents act directly (executing multi-step tasks, making routine decisions within defined bounds) paired with a human override/review layer that lets a person inspect, pause, redirect, or roll back what the agent is doing or about to do. Expect this to show up first as an “agent mode” bolted onto existing human UIs (a visible activity log, an approve/reject queue, a kill switch), rather than as a wholesale redesign — most products will run both surfaces side by side rather than picking one.
-
Confidence: Medium
-
Falsification: Mainstream productivity and enterprise software ships agent-driven automation without any corresponding human-visible override, audit trail, or pause/rollback mechanism, and users show no measurable demand for one — i.e., agent action happens invisibly rather than through a legible, interruptible surface.
#55. More Cyborgs (2-4 years)
Cyborg capabilities like neural implants and health trackers will get better and more mainstream. Some extreme high-performance seekers will advance the art, patching AI-powered glasses and other interfaces directly into their neurology.
-
Confidence: Low
-
Falsification: Neural implant adoption remains confined to narrow medical-necessity cases with no meaningful movement toward elective/mainstream use; health tracker capability improvement continues at pre-existing (non-accelerated) pace.
#56. AX Matures Into a First-Class Consideration (2-4 years)
The agent-interface/human-override pattern from #52 matures from a bolted-on afterthought into a designed-in architectural layer: permissioning, audit, and override become a standard part of the product spec from day one (the way accessibility or security review is supposed to be today), rather than something retrofitted after an agent-automation feature ships. Design systems and component libraries add agent-facing primitives (confirmation gates, explainability panels, reversible-action patterns) as a recognized category alongside buttons and forms.
-
Confidence: Medium
-
Falsification: Agent oversight/override tooling stays a bespoke, one-off feature built separately by each product team rather than converging into shared design-system primitives or platform conventions within the period.
Physical AI / Robotics
#57. Physical AI Remains an Exception (within 12 months)
Physical AI stays early and uneven: humanoid and mobile-robot demos multiply and warehouse/logistics pilots expand meaningfully, but general-purpose dexterous manipulation in unstructured environments (homes, general retail, unstructured outdoor work) remains firmly in pilot phase rather than reaching commercial-scale deployment.
-
Confidence: High
-
Falsification: General-purpose home or unstructured-retail robotic deployment reaches meaningful commercial scale (beyond narrow pilots or demos) within the period.
#58. Physical AI Finds Its Feet (3-5 years)
Physical AI reaches an inflection point comparable to where LLM-based coding agents are today: narrow, well-instrumented environments (warehouses, manufacturing lines, agriculture, controlled logistics) see real commercial-scale deployment, while general-purpose home/consumer robotics stays behind — gated by the same trust, liability, and absorption dynamics already seen in regulated software (see Regulated Industries below).
-
Confidence: Low
-
Falsification: Commercial-scale physical AI deployment fails to materialize even in narrow, controlled environments (warehouses, manufacturing) within the period, suggesting the constraint is technical rather than trust/liability/absorption-based, OR general-purpose home robotics reaches commercial scale ahead of narrow-environment deployment, inverting the predicted order.
Regulated Industries: Healthcare, Finance, Government
#59. AI Adoption Begins to Catch Up to Tech (2-4 years)
AI adoption in regulated industries (healthcare, finance, government) lags the general-enterprise pod/pace predictions above by 2-3 years, not because the technology isn’t ready, but because compliance, audit, and liability requirements gate rollout speed independent of technical capability — meaning the “half the people, quarter of the time” pod dynamic shows up in these sectors’ back-office and internal tooling work well before it shows up in patient-facing, transaction-facing, or citizen-facing systems.
-
Confidence: High
-
Falsification: Regulated-industry, customer-facing AI deployment speed converges with general-enterprise deployment speed within the period rather than lagging it.
A note on how to read this: confidence levels here reflect my sense of variance in outcome, not how much I want each prediction to be true. Where two of my own predictions are in tension, I’ve tried to flag it inline rather than paper over it — most notably the customer-absorption-speed constraint (Org Structure & Delivery Model, #2) against the sustained multi-fold delivery speedups predicted: if absorption really is the bottleneck, most of the pod-speed gain shows up as freed internal capacity (tech-debt burn-down, per #1) rather than a sustained 8x increase in customer-facing throughput. Similarly, #3 (managers won’t take on more reports) and #4 (pods shrink to 2-3 people) collide on org-chart math if headcount doesn’t grow elsewhere — I think more pods per manager is the likelier resolution, but it’s an open question, not a settled one. I’ll be revisiting the higher-uncertainty predictions in this piece over the next year as evidence comes in.
What do you think? What did I get right, what did I miss? What are you seeing that supports or pushes back on these?
_Keith MacKay is a technology strategy consultant and CTO in EY-Parthenon’s Software Strategy Group (SSG), specializing in AI disruption and technology diligence for private equity and corporate clients. SSG’s AI Disruption Lab conducts rapid assessments of how AI transforms and threatens existing business models and value chains. Keith teaches at Northeastern University and writes about strategy, management, and AI/technology, with Claude and Codex as AI collaborators. For more, visit his Substack or Medium site.