FP8 compute and you
Primer on FP8 and the iPhone 18 Pro / A20 Pro
September 17, 2026
Apple's new iPhone 18 Pro is out tomorrow, and I am eagerly awaiting mine. Eva and I are still on launch iPhone 13 Pros, which are generally working fine after 5 years, even on their original batteries (79 and 82% health respectively). This is my first 5 year phone, and I suspect the 18 Pro will be the next.
A big part of this is that Apple added native FP8 support to the A20 Pro (and related M6) SoC. I see a lot of people missing how exciting this is in the online discourse, so here is my little blurb on FP8 compute.
Floating point numbers are how computers store fractional numbers. This in contrast to integers. The size of an integer is usually written like INT32 or INT64, with the number of bits represented telling you how big a number you can store. Floating point numbers like FP32 (standard precision) or FP64 (double precision) generally concern themselves with how precise a number they can store.
Before the AI craze the big differentiator between consumer GPUs and enterprise GPUs was the ability to execute FP64 numbers at full speed. Scientific compute required the extra precision for research.
FP16 is a change in the opposite direction. Lots of use cases of floating numbers do not need the full precision of a 32-bit number. This means quite a lot of memory saved and faster compute when GPUs can execute two of these operations at a time. Amusingly, I think NVIDIA first dabbled with native FP16 on the Tegra X1 (the processor that ended up in the Nintendo Switch) - likely to save VRAM. Eventually both AMD and NVIDIA adopted FP16 across their datacenter and consumer lines broadly.
Importantly FP16 and its cousin BF16 (without going too deep, a more efficient way to capture the dynamic range of FP32 with only 16 bits of data) became a massive performance unlock for LLM models. BF16 allowed for full fidelity with 2x compute uplift and 1/2 the memory usage on NVIDIA's Ampere datacenter GPUs and Ada Lovelace consumer GPUs (A100 and 40x0 consumer GPUs).
Fast forwarding a bit (because otherwise this is going to be less of a primer) but researchers have realized that even lower precision numbers allows for negligible quality loss (often less than 0.5%) and further memory savings and performance gains. Meet FP8! Supported natively on NVIDIA's Hopper datacenter and Ada Lovelace consumer GPUs (H100, 40x0, etc).
Early usage of FP8 was often post training - a model was produced at higher precision and quantized down to FP8 to perform inference for end-users. However the folks at Deepseek have even trained models in FP8 significantly reducing the reliance on compute resources. FP8 is a sweet spot, essentially perfect performance with much less memory used and that brings us back to the A20 Pro and M6 chips.
Apple is going to build its models around FP8 (and mixed precision, more on that at the end) and take the memory savings and performance uplift. These models will still see the memory savings on older Apple chips (they support their AI models on all M series processors and the A17 Pro and later) but without the 2x compute speed uplift. These FP8 native phones will surely be faster and for longer as Apple continues to develop more on-device AI models.
So that's it. I am excited for more on device AI, and the A20 Pro is likely going to be a monster for years to come.
Bonus: Mixed precision! So the "final frontier" in inference is mixed precision floating point numbers. MXFP4 and NVFP4 are two competing formats, the latter being NVIDIA's proprietary format. By performing fast native FP4 math and accumulating into larger FP16/BF16/FP32 registers we see yet another massive savings in memory and compute uplift. These mixed precision quantization strategies do risk more quality loss, but with modern reasoning models it is less of an issue in terms of how effective a model is at solving a task. I do expect Apple to hop on the native FP4 compute train eventually, but for the kinds of models they are deploying on-device it may be a lot less useful than for big sparse reasoning models.