Apple M3 Ultra GPU (80-core) vs Max GPU (40-core)

We performed a head-to-head comparison of the Apple M3 Ultra GPU (80-core) with 1280 pipelines and 10240 shaders against the 1 year and 4 months older Max GPU (40-core) that utilizes 640 pipelines and 5120 shaders. Here you will find complete details about specs, efficiency, performance tests, and more.

Review

General comparison of performance in games, applications, power efficiency, and other metrics
Gaming
Performance in DirectX, OpenCL, and Vulkan games
Workstation
Perf. in 3D modeling, video editing and rendering apps
AI/ML
Capabilities for machine learning and AI-related tasks
Energy Efficiency
Power consumption efficiency in different scenarios
NanoReview Final Score
Overall video card score
The "Energy Efficiency" metric has less impact on the NanoReview Score for desktop GPU.

Key differences

Key distinctions and advantages of M3 Max GPU (40-core) over M3 Ultra GPU (80-core)
Reasons to consider the Apple M3 Ultra GPU (80-core)
  • Performs significantly better (up to 58%) in 3DMark Steel Nomad Lite
  • 2x higher maximum theoretical performance (28.3 vs 14.1 TFLOPS)
  • Has 2x higher memory bandwidth: 819.3 vs 409.6 GB/s
  • Achieves 56% more points in the GeekBench 6 Compute test (147K vs 94K)
  • Has 2x more shading units (10240 vs 5120)

Benchmarks

Graphics cards’ performance in recent benchmarking apps

3D Mark

Multiplatform graphics benchmark suite that directly correlates with performance in modern games
Steel Nomad Lite Score
Solar Bay 63597 49656
Wild Life Extreme 47814 30893
Sources: 3DMark [1], [2]

GeekBench 6 OpenCL

GPU test for computational tasks (image processing, photography, computer vision, and ML)
GB6 Compute Score
Background Blur 204.5 img/sec 157.6 img/sec
Face Detection 126 img/sec 100.9 img/sec
Horizon Detection 6.53 Gpixels/sec 3.99 Gpixels/sec
Edge Detection 10.7 Gpixels/sec 6.04 Gpixels/sec
Gaussian Blur 9.5 Gpixels/sec 5 Gpixels/sec
Feature Matching 1.05 Gpixels/sec 0.84 Gpixels/sec
Stereo Matching 560.3 Gpixels/sec 363.8 Gpixels/sec
Particle Physics 25267.8 FPS 12226.9 FPS
API OpenCL OpenCL
Sources: Geekbench [3], [4]

Cinebench 2024 GPU

Hardware benchmark using Maxon's Cinema 4D rendering engine
Cinebench 2024 GPU

Blender

Rendering performance test for 3D modeling
Blender GPU
Sources: Blender [9], [10] – 31 & 446 samples

Recent User Tests

The latest benchmark tests that have been submitted by users
Apple M3 Ultra GPU (80-core)
DateBenchmarkResult
📘 2025-03-16 (Tyler)Geekbench 6 OpenCL145706
📘 2025-03-16 (Tyler)Cinebench 202419448
Apple M3 Max GPU (40-core)
No benchmark results yet

Artificial Intelligence Tests

Performance in machine learning and artificial intelligence tasks

GeekBench 6 ML

Tests throughput of AI operations in single, half, and quantized precision
GB6 ML Single Precision
GB6 ML Half Precision
GB6 ML Quantized
Image Classification (SP) 8094 8685
Image Segmentation (HP) 24595 20739
Image Super Resolution (Q) 25190 22775
Face Detection (HP) 38337 37676
Pose Estimation (Q) 111021 91609
Text Classification (SP) 2786 2909
Machine Translation (HP) 4339 5274
Object Detection (SP) 7668 7480
Depth Estimation (Q) 37488 38212
Style Transfer (SP) 244483 190681
Framework Core ML Core ML
Backend GPU GPU
Sources: Geekbench [9], [10]

Specifications

Technical specifications of Apple M3 Ultra GPU (80-core) and Max GPU (40-core)

General

Vendor Apple Apple
Build Integrated Integrated
Released March 5, 2025 October 31, 2023
Case Desktop Laptop
Purpose Professional Professional
Segment High-end High-end
Architecture Apple M GPU Apple M GPU
GPU Codename Custom -
Rival Equivalent - GeForce RTX 5070 Ti - GeForce RTX 4070 Laptop
Successor - Apple M4 Ultra GPU (80-core) - Apple M5 Max GPU (40-core)
Recommended CPU - Apple M3 Ultra or above - Apple M3 Max or above
Used in CPUs - Apple M3 Ultra - Apple M3 Max
Desktop GPU rating (#29th place)
Laptop GPU ranking (#21st place)

Graphics Processing Unit

Base Clock - 500 MHz
Boost Clock 1380 MHz 1380 MHz
Shading Units 10240 5120
Texture Mapping Units (TMUs) 640 320
Render Output Units (ROPs) 320 160
Compute Units (Pipelines) 1280 640
Instructions Per Cycle 2 IPC 2 IPC

Raw Performance

Pixel Fill Rate 442 GPixel/s 221 GPixel/s
Texture Fill Rate 883 GTexel/s 442 GTexel/s
FLOPS (FP32)
28.3 TFLOPS
14.1 TFLOPS

Physical

Interface Custom Custom
TGP 140 W 60 W
Manufacturing TSMC TSMC
Fabrication Process 3 nm 3 nm
Transistor Count 184 billion 56 billion
Max. Temperature 100°C 100°C

Memory

Memory Type System Shared System Shared
Memory Clock 6400 MHz 6400 MHz
Effective Memory Speed - 12800 Mbps
Bus 1024-bit 512-bit
ECC No No
Memory Bandwidth
819.3 GB/s

API

Ray Tracing Yes Yes
DLSS No No

Cast your vote

Choose between two graphics cards
2 (100%)
0 (0%)
Total votes: 2

User opinions

You can share your opinion or ask a question in the comments below
🌐 Register your profile and become part of NanoReview community!