Apple M4 Max GPU (40-core) vs M3 Pro GPU (18-core)

We compared two integrated laptop professional GPUs: the Apple M4 Max GPU (40-core) with 640 pipelines and 5120 shaders against the 1 year older M3 Pro GPU (18-core) that utilizes 288 pipelines and 2304 shaders. Here you will find complete details about specs, efficiency, performance tests, and more.

Review

General comparison of performance in games, applications, power efficiency, and other metrics
Gaming
Performance in DirectX, OpenCL, and Vulkan games
Workstation
Perf. in 3D modeling, video editing and rendering apps
AI/ML
Capabilities for machine learning and AI-related tasks
Energy Efficiency
Power consumption efficiency in different scenarios
NanoReview Final Score
Overall video card score

Key differences

Key distinctions and advantages of M3 Pro GPU (18-core) over M4 Max GPU (40-core)
Reasons to consider the Apple M4 Max GPU (40-core)
  • Performs significantly better (up to 2.4x) in 3DMark Steel Nomad Lite
  • 2.5x higher maximum theoretical performance (16.2 vs 6.4 TFLOPS)
  • Has 3.6x higher memory bandwidth: 546 vs 153.6 GB/s
  • Achieves 2.7x more points in the GeekBench 6 Compute test (116K vs 43K)
  • Has 2.2x more shading units (5120 vs 2304)

Benchmarks

Graphics cards’ performance in recent benchmarking apps

3D Mark

Multiplatform graphics benchmark suite that directly correlates with performance in modern games
Steel Nomad Lite Score
Solar Bay 61312 22637
Wild Life Extreme 36641 14021
Sources: 3DMark [1], [2]

GeekBench 6 OpenCL

GPU test for computational tasks (image processing, photography, computer vision, and ML)
GB6 Compute Score
Background Blur 195 img/sec 82.1 img/sec
Face Detection 125.9 img/sec 53.6 img/sec
Horizon Detection 4.82 Gpixels/sec 1.82 Gpixels/sec
Edge Detection 7.47 Gpixels/sec 2.59 Gpixels/sec
Gaussian Blur 5.8 Gpixels/sec 1.84 Gpixels/sec
Feature Matching 1.06 Gpixels/sec 0.48 Gpixels/sec
Stereo Matching 417.4 Gpixels/sec 148.6 Gpixels/sec
Particle Physics 17079.9 FPS 4893.9 FPS
API OpenCL OpenCL
Sources: Geekbench [3], [4]

Cinebench 2024 GPU

Hardware benchmark using Maxon's Cinema 4D rendering engine
Cinebench 2024 GPU

Blender

Rendering performance test for 3D modeling
Blender GPU
Sources: Blender [9], [10]456 & 237 samples

Recent User Tests

The latest benchmark tests that have been submitted by users
Apple M4 Max GPU (40-core)
DateBenchmarkResult
📘 2025-07-10 (thomas)Geekbench 6 OpenCL116478
Apple M3 Pro GPU (18-core)
No benchmark results yet

Artificial Intelligence Tests

Performance in machine learning and artificial intelligence tasks

GeekBench 6 ML

Tests throughput of AI operations in single, half, and quantized precision
GB6 ML Single Precision
GB6 ML Half Precision
GB6 ML Quantized
Image Classification (SP) 10302 5578
Image Segmentation (HP) 23763 9924
Image Super Resolution (Q) 29314 13323
Face Detection (HP) 42016 19900
Pose Estimation (Q) 105972 35730
Text Classification (SP) 2963 2904
Machine Translation (HP) 6594 4985
Object Detection (SP) 8935 5496
Depth Estimation (Q) 43143 23344
Style Transfer (SP) 221563 78342
Framework Core ML Core ML
Backend GPU GPU
Sources: Geekbench [9], [10]

Specifications

Technical specifications of Apple M4 Max GPU (40-core) and M3 Pro GPU (18-core)

General

Vendor Apple Apple
Build Integrated Integrated
Released October 30, 2024 October 31, 2023
Case Laptop Laptop
Purpose Professional Professional
Segment Mid-range Mid-range
Architecture Apple M GPU Apple M GPU
GPU Codename Custom Custom
Rival Equivalent - GeForce RTX 4070 Laptop - GeForce RTX 4060 Laptop
Successor - Apple M5 Max GPU (40-core) - Apple M5 Pro GPU (20-core)
Recommended CPU - Apple M4 Max (16-Core) or above - Apple M3 Pro or above
Used in CPUs - Apple M4 Max (16-Core) - Apple M3 Pro
Laptop GPU ranking (14th and 60th place)

Graphics Processing Unit

Base Clock 500 MHz 500 MHz
Boost Clock 1578 MHz 1380 MHz
Shading Units 5120 2304
Texture Mapping Units (TMUs) 320 144
Render Output Units (ROPs) 160 72
Compute Units (Pipelines) 640 288
Instructions Per Cycle 2 IPC 2 IPC

Raw Performance

Pixel Fill Rate 252 GPixel/s 99 GPixel/s
Texture Fill Rate 505 GTexel/s 199 GTexel/s
FLOPS (FP32)
16.2 TFLOPS

Physical

Interface Custom Custom
TGP 50 W 30 W
Manufacturing TSMC TSMC
Fabrication Process 3 nm 3 nm
Transistor Count - 25.2 billion
Max. Temperature 100°C 100°C

Memory

Memory Type System Shared System Shared
Memory Clock 8533 MHz 6400 MHz
Effective Memory Speed - 12800 Mbps
Bus 512-bit 192-bit
ECC No No
Memory Bandwidth
546 GB/s

API

Ray Tracing Yes Yes
DLSS No No

Cast your vote

Choose between two graphics cards
1 (100%)
0 (0%)
Total votes: 1

User opinions

You can share your opinion or ask a question in the comments below
🌐 Register your profile and become part of NanoReview community!