iPhone 18 Pro’s A20 Pro Chip Doubles AI Power [2026]

Apple used its September 9 event to introduce the iPhone 18 Pro and iPhone 18 Pro Max, and the headline change is not the camera or the screen. It is the chip. The new A20 Pro ships with a dual 16-core Neural Engine, effectively doubling the on-device AI compute available to Apple Intelligence compared with last year’s A19 Pro, according to Apple’s newsroom announcement. The move puts Apple’s on-device AI hardware directly against Samsung’s Snapdragon 8 Elite Gen 5 inside the Galaxy S26 Ultra and Google’s Tensor G5 inside the Pixel 10 Pro, and it reframes how the three biggest phone makers are fighting over the same real estate: who can run the most capable AI model without sending your data to a server.

This is not a minor spec bump. For three straight generations, from the A17 Pro through the A19 Pro, Apple’s Neural Engine stayed locked at 16 cores and roughly 35 trillion operations per second, according to benchmark data first reported by MacRumors. The A20 Pro breaks that pattern by borrowing the dual-block Neural Engine architecture Apple debuted in the M6 chip for Mac and iPad, letting both 16-core blocks run in parallel. That is the kind of architectural shift Apple usually reserves for a multi-year roadmap, not a routine annual refresh.

Google · Preferred Sources

Don't miss new tech stories on Google

Add Tech Insider once in the Google app and our stories appear in your news suggestions.

Add Now

What Apple Announced: An AI-First iPhone 18 Pro

Apple’s September 9 keynote, covered live by CNBC, framed the iPhone 18 Pro and iPhone 18 Pro Max around Apple Intelligence running natively on the device rather than routing requests to a data center. The pitch is simple: faster responses, no network dependency, and data that never leaves the phone. Tech Insider covered the pricing and hardware side of the launch separately, including the $1,199 starting price and foldable reveal, but the AI architecture underneath deserves its own look because it is the part of the announcement Apple spent the most stage time defending.

The A20 Pro pairs a six-core CPU, which Apple says runs up to 20% faster than the previous generation’s Pro chip, with neural accelerators built directly into the CPU cores. That detail matters because it means AI inference is no longer confined to the Neural Engine block alone. Low-latency tasks, like predicting the next word in a sentence or classifying a photo as it is captured, can now run on whichever silicon block is fastest for that specific job. Apple has used a similar approach on the Mac side with the M7 chip, part of the company’s broader AI-first silicon roadmap built around the Baltra architecture.

Inside the A20 Pro: What the Dual Neural Engine Actually Does

A Neural Engine is a dedicated block of silicon built to run the matrix math behind machine learning models faster and more efficiently than a general-purpose CPU or GPU. Apple has shipped one in every iPhone since the A11 Bionic in 2017, but the core count and throughput barely moved for years. The jump to a dual 16-core design, for a total of 32 AI-dedicated cores, is Apple’s answer to a problem every phone maker now shares: modern on-device AI models are bigger and hungrier than the chips built to run them.

System frameworks can now address both 16-core blocks at once, which Apple says roughly doubles peak throughput versus a single-block design. In practice, that headroom goes toward two things: running larger language models locally, and running more AI features simultaneously without one task stalling another. On the A18 Pro, Apple’s Neural Engine could handle up to eight models running in parallel, up from six on the A17 Pro. The A20 Pro is built to push that concurrency further, which is what allows Siri, camera AI, and background text suggestions to all run at once without a visible slowdown.

The Neural Engine’s Five-Generation Sprint

Looking at the last five generations shows how sharply Apple has accelerated its AI silicon roadmap after years of incremental gains. The table below tracks publicly reported Neural Engine specifications across recent A-series chips, based on Apple’s own chip pages and benchmark data reported by MacRumors.

ChipNeural EnginePeak ThroughputParallel ModelsOn-Device LLM
A16 Bionic16-core~17 TOPS4-6Not supported
A17 Pro16-core~35 TOPS6Not supported
A18 Pro16-core~35 TOPS83B parameters
A19 Pro16-core~35 TOPS83B+ parameters
A20 ProDual 16-core (32 total)~2x A19 Pro8+Larger on-device model, size unconfirmed

Apple has not published an exact TOPS figure or parameter count for the A20 Pro’s on-device model as of this writing, so treat the “2x A19 Pro” figure as Apple’s own characterization rather than an independently verified benchmark. What is clear is the trend line: after three years of a flat 35 TOPS ceiling, Apple chose to double the physical Neural Engine footprint rather than simply clock it higher, a sign that raw efficiency gains from process shrinks alone were no longer enough to keep pace with larger on-device models.

Siri’s Overhaul: What Changes for the Assistant

Apple has tied the A20 Pro’s extra Neural Engine capacity directly to an upgraded Siri experience built on Apple Intelligence, according to the company’s own hardware materials. The framing is that Siri should understand context and generate responses locally rather than calling out to a server for every request. That has been Apple’s stated goal since Apple Intelligence launched, but the company has repeatedly slipped on shipping the more ambitious, personalized version of Siri. The A20 Pro’s extra compute headroom removes one of the technical excuses, since a bigger on-device model needs a bigger Neural Engine to run it without draining the battery or introducing lag.

Whether Apple actually ships the deeper Siri overhaul on this hardware, or pushes it to a future iOS update, remains the open question analysts are watching. Apple has a track record of announcing AI silicon well ahead of the software features that fully exploit it, so the safest read is that the A20 Pro sets the ceiling for what is possible, not a guarantee of what ships on day one.

How On-Device LLMs Actually Run Without the Cloud

Apple Intelligence’s on-device language model has run at roughly 3 billion parameters since the A18 Pro generation, generating text at around 30 tokens per second locally. That is small compared with cloud models like GPT-6 Astra or Gemini 3.8 Flash, both of which run on server clusters with none of a phone’s power or thermal limits. Tech Insider’s head-to-head comparison of GPT-6 Astra, Claude Opus 5, and Gemini 3.8 Flash shows just how far ahead cloud models remain on raw benchmark scores. The tradeoff Apple is betting on is that most everyday phone tasks, rewriting a text, summarizing a notification, tagging a photo, do not need a frontier-scale model. They need a fast, private, always-available one.

Apple has not confirmed the exact parameter count of the A20 Pro’s on-device model, so any specific figure beyond “larger than the A19 Pro’s” should be treated as unverified until Apple’s developer documentation updates. Historically, Apple has revealed these details months after launch through WWDC sessions and Core ML updates rather than at the September keynote itself.

iPhone 18 Pro vs Galaxy S26 Ultra vs Pixel 10 Pro: The AI Silicon Race

Apple is not the only company betting its flagship phone on local AI silicon. Samsung’s Galaxy S26 Ultra runs on the Snapdragon 8 Elite Gen 5 for Galaxy, whose Hexagon NPU Samsung says is 39% faster than the chip in the previous Galaxy generation, alongside a 19% faster CPU and 24% faster GPU, according to Samsung’s own announcement. Google’s Pixel 10 Pro uses the Tensor G5, whose AI-focused TPU block is roughly 60% faster than the Tensor G4’s, based on figures Google has published for its Gemini Intelligence features.

PhoneChipAI BlockGen-over-Gen AI GainCore AI Assistant
iPhone 18 Pro / Pro MaxA20 ProDual 16-core Neural Engine~2x (Apple’s figure)Siri via Apple Intelligence
Galaxy S26 UltraSnapdragon 8 Elite Gen 5 for GalaxyHexagon NPU+39% NPUGalaxy AI + Gemini Intelligence
Pixel 10 ProTensor G5Custom TPU+60% TPUGemini Intelligence (Nano v3)

The three approaches differ in philosophy as much as raw numbers. Apple is doubling down on privacy-first, fully local processing and tightly coupling it to Siri. Samsung leans on Qualcomm’s raw NPU throughput to power camera-heavy Galaxy AI features like generative edit and AI zoom. Google splits the difference, running lighter tasks locally through Gemini Nano v3 while routing anything demanding to the cloud-based Gemini Intelligence stack, a strategy Tech Insider covered in its look at Qualcomm’s own dual Snapdragon 8 Elite Gen 6 chip announcement aimed at the same on-device AI race.

Samsung’s Bet: Raw NPU Power for Camera-Centric AI

Samsung’s strategy with the Galaxy S26 Ultra leans hardest into media processing. The Hexagon NPU inside the Snapdragon 8 Elite Gen 5 for Galaxy is tuned for high-resolution image and video AI work: generative photo editing, AI-assisted zoom, and real-time video enhancement. That focus reflects where Samsung believes most consumers actually use AI on a phone, in the camera app, rather than in a chat-style assistant. Samsung also ties Galaxy AI to Google’s Gemini Intelligence stack for devices that support Gemini Nano v3, giving Galaxy S26 Ultra owners a hybrid setup: fast local NPU processing for camera tasks, and cloud-backed Gemini for anything requiring deeper reasoning.

Google’s Bet: Gemini Intelligence and Always-On Assistance

Google’s Tensor G5 takes a different path. Rather than chasing the highest raw TOPS figure, Google has optimized the Pixel 10 Pro’s TPU for low-power, always-on tasks: live call screening, real-time captioning, and background summarization that runs continuously without draining the battery. Full Gemini Intelligence, Google’s more capable on-device and cloud-hybrid assistant, requires Gemini Nano v3 support, currently limited to the Pixel 10 series and Samsung’s Galaxy S26 lineup, according to Google’s own Gemini product blog. In multiple 2026 comparisons, reviewers have called the Pixel 10 Pro the strongest overall “AI phone” specifically because of how deeply Gemini Intelligence integrates with Google Workspace and search, not because of raw chip specs.

Why On-Device AI Matters Now: Privacy, Latency, and Cost

Three forces are pushing all three companies toward local AI silicon at the same time. Privacy concerns around cloud AI have grown louder as regulators in the EU and several US states scrutinize how AI companies handle personal data. Running inference on-device sidesteps a lot of that exposure because the data never leaves the phone. Latency is the second factor. A cloud round trip adds anywhere from a few hundred milliseconds to a few seconds depending on network conditions, which breaks the illusion of a responsive assistant. And cost is the third, and arguably the most underrated, factor: every cloud AI query costs the company running it real money in compute. Shifting inference onto the customer’s own chip is a direct way to cut operating expenses at scale, something that matters enormously once hundreds of millions of phones are making AI requests every day.

Market Impact: What This Means for the Smartphone AI Race

Apple’s move raises the floor for what counts as an “AI flagship” heading into 2027. If the A20 Pro’s dual Neural Engine performs as advertised, every major Android chipmaker, not just Qualcomm and Google’s in-house silicon, will face pressure to match or beat that architecture within a generation or two. That competitive pressure typically shows up fastest in the mid-range, where chipmakers trickle down flagship AI features to keep volume segments competitive. Expect Qualcomm’s next Snapdragon 8-series chip and MediaTek’s Dimensity flagships to lean harder into dual-block or multi-block NPU designs by late 2026 or early 2027.

There is also a supply chain angle. Bigger Neural Engines and neural accelerators built into the CPU require more silicon area, which increases die size and cost per chip. Apple has already been managing tighter memory supply through long-term deals, including the S11 chip inside the Apple Watch Series 12 and broader NAND sourcing agreements, so a bigger, more AI-dense A20 Pro adds one more input cost Apple has to absorb or pass on to consumers.

The Developer Angle: Core ML and Third-Party AI Apps

The Neural Engine upgrade is not just about Apple’s own features. Third-party developers access it through Apple’s Core ML framework, which lets apps run their own machine learning models on-device rather than calling an external API. A doubled Neural Engine means developers can ship larger, more capable models inside their apps, things like on-device translation, offline image generation, or local voice transcription, without needing a network connection or paying for cloud inference themselves. Historically, Core ML performance gains have shown up clearly in benchmarks: Apple’s own testing during the iOS 18 cycle showed the A17 Pro’s Core ML Neural Inference Score jumping from 6,249 to 7,816 on Geekbench after a software update alone, before any new hardware shipped. That kind of software-plus-hardware compounding is exactly what Apple is betting on again with the A20 Pro.

Predictions: Where On-Device AI Goes From Here

  • Apple will likely disclose the A20 Pro’s exact on-device model size and TOPS figure at WWDC 2027, following its pattern of revealing deeper technical specs months after the September keynote.
  • Qualcomm and MediaTek will both move toward dual-block or multi-block NPU designs within the next one to two flagship chip cycles to close the gap with the A20 Pro’s architecture.
  • Expect Siri’s more personalized, context-aware features, previously delayed multiple times, to ship in stages through 2027 rather than all at once, as Apple validates the A20 Pro’s real-world performance.
  • Mid-range phones from all three ecosystems will start advertising “on-device AI” as a headline spec by 2027, following the same trickle-down pattern seen with 5G and high-refresh displays.
  • Cloud AI providers will increasingly position their models as complements to on-device AI rather than replacements, since phone makers have made clear that local processing is now a permanent part of the stack, not a stopgap.

Historical Context: How Apple Got Here

Apple’s Neural Engine debuted quietly in the A11 Bionic back in 2017, mostly to power Face ID and Animoji. It took years for the block to become a marketing centerpiece. The real turning point came with Apple Intelligence’s 2024 launch, which forced Apple to justify years of Neural Engine investment as the foundation for a genuine on-device AI assistant rather than a background feature. The three-year plateau at 35 TOPS across the A17 Pro, A18 Pro, and A19 Pro now reads as a transitional period, one where Apple’s software ambitions for Apple Intelligence outpaced what the existing Neural Engine architecture could comfortably support. The A20 Pro’s dual-block redesign, borrowed directly from the M6 chip built for Mac and iPad, is Apple’s clearest signal yet that it is willing to redesign core silicon architecture specifically to keep pace with AI demands, rather than treating AI as a feature bolted onto an existing chip roadmap.

What Buyers Should Actually Expect

For everyday users, the practical impact of the A20 Pro’s Neural Engine will show up gradually rather than all at once. Camera AI features, live translation, and on-device text generation should feel noticeably snappier at launch. The more ambitious Siri overhaul Apple has promised, and repeatedly delayed, is the feature most likely to arrive later through software updates rather than being fully available on day one. Buyers comparing the iPhone 18 Pro against the Galaxy S26 Ultra or Pixel 10 Pro should weigh what kind of AI experience they actually want: Apple’s privacy-first, fully local approach, Samsung’s camera-focused NPU power, or Google’s cloud-hybrid Gemini Intelligence, since each company is optimizing its silicon for a genuinely different use case rather than simply chasing the same benchmark.

Frequently Asked Questions

What is the A20 Pro chip?
The A20 Pro is Apple’s newest flagship chip, introduced with the iPhone 18 Pro and iPhone 18 Pro Max on September 9, 2026. It features a dual 16-core Neural Engine, for 32 total AI-dedicated cores, and a six-core CPU with built-in neural accelerators.

How much faster is the A20 Pro’s AI performance than the A19 Pro?
Apple says the dual Neural Engine design roughly doubles peak AI compute compared with the A19 Pro’s single 16-core Neural Engine. Apple has not published independent third-party benchmark figures confirming this as of this writing.

Does the iPhone 18 Pro run AI models fully on-device?
Apple Intelligence’s core assistant features run locally on the A20 Pro’s Neural Engine, following the same on-device approach used since the A18 Pro generation. More complex tasks may still use Apple’s cloud infrastructure, called Private Cloud Compute, depending on the request.

How does the iPhone 18 Pro’s AI compare to the Galaxy S26 Ultra?
The Galaxy S26 Ultra uses Samsung’s Snapdragon 8 Elite Gen 5 for Galaxy, with a Hexagon NPU that Samsung says is 39% faster than the prior generation. It leans toward camera and media-focused AI features, while the iPhone 18 Pro emphasizes assistant-style, fully local processing through Siri.

How does the iPhone 18 Pro’s AI compare to the Pixel 10 Pro?
The Pixel 10 Pro runs Google’s Tensor G5 chip, with a TPU block roughly 60% faster than the Tensor G4’s. It powers Gemini Intelligence, which blends on-device Gemini Nano v3 processing with cloud-based Gemini for more demanding tasks, a more hybrid approach than Apple’s local-first strategy.

Will the new Siri features work on older iPhones?
Apple has not confirmed which Siri features will be exclusive to the A20 Pro’s dual Neural Engine versus available on older Apple Intelligence-compatible chips. Historically, Apple has reserved its most demanding on-device AI features for the newest silicon generation.

When will Apple reveal more technical details about the A20 Pro’s AI model?
Apple typically shares deeper Core ML and on-device model specifications at its Worldwide Developers Conference, which would put a fuller technical disclosure around WWDC 2027 rather than at the September hardware launch.

Is on-device AI actually better than cloud AI?
Neither approach is strictly better. On-device AI offers faster response times, works offline, and keeps data local, but runs smaller models with less raw capability. Cloud AI models like those compared in Tech Insider’s GPT-6 Astra, Claude Opus 5, and Gemini 3.8 Flash benchmark handle far more complex reasoning but require a network connection and raise more data privacy questions.

Related Coverage

Nadia Dubois

Nadia Dubois

AI & Innovation Editor

Nadia Dubois is the AI & Innovation Editor at Tech Insider, where she tracks the rapid evolution of artificial intelligence, from foundation models to real-world enterprise deployment. She previously covered AI and startups for La Tribune and contributed to MIT Technology Review's European coverage. Nadia specializes in generative AI, AI regulation, and the intersection of technology and European industrial policy. She holds a dual degree in Computational Linguistics and Journalism from Sciences Po Paris.

View all articles