Apple M3 Max GPU (40-core) vs M1 Pro GPU (16-core)

We compared two integrated laptop professional GPUs: the Apple M3 Max GPU (40-core) with 640 pipelines and 5120 shaders against the 2 years and 1 month older M1 Pro GPU (16-core) that utilizes 256 pipelines and 2048 shaders. Here you will find complete details about specs, efficiency, performance tests, and more.

Review

General comparison of performance in games, applications, power efficiency, and other metrics
Gaming
Performance in DirectX, OpenCL, and Vulkan games
Workstation
Perf. in 3D modeling, video editing and rendering apps
AI/ML
Capabilities for machine learning and AI-related tasks
Energy Efficiency
Power consumption efficiency in different scenarios
NanoReview Final Score
Overall video card score

Key differences

Key distinctions and advantages of M1 Pro GPU (16-core) over M3 Max GPU (40-core)
Reasons to consider the Apple M3 Max GPU (40-core)
  • Performs significantly better (up to 3.1x) in 3DMark Steel Nomad Lite
  • 2.7x higher maximum theoretical performance (14.1 vs 5.3 TFLOPS)
  • Manufactured using a more efficient 3 nm process technology
  • Achieves 2.2x more points in the GeekBench 6 Compute test (94K vs 42K)
  • Has 2x higher memory bandwidth: 409.6 vs 204.8 GB/s
  • Has 2.5x more shading units (5120 vs 2048)

Benchmarks

Graphics cardsโ€™ performance in recent benchmarking apps

3D Mark

Multiplatform graphics benchmark suite that directly correlates with performance in modern games
Steel Nomad Lite Score
Solar Bay 49656 12485
Wild Life Extreme 30893 9873
Sources: 3DMark [1], [2]

GeekBench 6 OpenCL

GPU test for computational tasks (image processing, photography, computer vision, and ML)
GB6 Compute Score
Background Blur 157.6 img/sec 74.5 img/sec
Face Detection 100.9 img/sec 47.3 img/sec
Horizon Detection 3.99 Gpixels/sec 1.73 Gpixels/sec
Edge Detection 6.04 Gpixels/sec 3.32 Gpixels/sec
Gaussian Blur 5 Gpixels/sec 1.75 Gpixels/sec
Feature Matching 0.84 Gpixels/sec 0.47 Gpixels/sec
Stereo Matching 363.8 Gpixels/sec 131.2 Gpixels/sec
Particle Physics 12226.9 FPS 5143 FPS
API OpenCL OpenCL
Sources: Geekbench [3], [4]

Cinebench 2024 GPU

Hardware benchmark using Maxon's Cinema 4D rendering engine
Cinebench 2024 GPU

Blender

Rendering performance test for 3D modeling
Blender GPU
Sources: Blender [9], [10] โ€“ 444 & 499 samples

Recent User Tests

The latest benchmark tests that have been submitted by users
Apple M3 Max GPU (40-core)
No benchmark results yet
Apple M1 Pro GPU (16-core)
DateBenchmarkResult
๐Ÿ“˜ 2026-03-27 (bbffx)Geekbench 6 OpenCL43467
๐Ÿ“˜ 2026-03-27 (bbffx)Cinebench 20242423
๐Ÿ“˜ 2026-03-27 (bbffx)Steel Nomad Light4048
๐Ÿ“˜ 2025-02-21 (serhii)Geekbench 6 OpenCL40932

Artificial Intelligence Tests

Performance in machine learning and artificial intelligence tasks

GeekBench 6 ML

Tests throughput of AI operations in single, half, and quantized precision
GB6 ML Single Precision
GB6 ML Half Precision
GB6 ML Quantized
Image Classification (SP) 8685 4112
Image Segmentation (HP) 20739 7488
Image Super Resolution (Q) 22775 9882
Face Detection (HP) 37676 14702
Pose Estimation (Q) 91609 33560
Text Classification (SP) 2909 2289
Machine Translation (HP) 5274 1629
Object Detection (SP) 7480 3633
Depth Estimation (Q) 38212 16986
Style Transfer (SP) 190681 53586
Framework Core ML Core ML
Backend GPU GPU
Sources: Geekbench [9], [10]

Specifications

Technical specifications of Apple M3 Max GPU (40-core) and M1 Pro GPU (16-core)

General

Vendor Apple Apple
Build Integrated Integrated
Released October 31, 2023 October 18, 2021
Case Laptop Laptop
Purpose Professional Professional
Segment High-end Mid-range
Architecture Apple M GPU Apple M GPU
Rival Equivalent - GeForce RTX 4070 Laptop - GeForce RTX 3050 Laptop
Successor - Apple M5 Max GPU (40-core) - Apple M5 Pro GPU (20-core)
Recommended CPU - Apple M3 Max or above - Apple M1 Pro or above
Used in CPUs - Apple M3 Max - Apple M1 Pro
Laptop GPU ranking (21st and 78th place)

Graphics Processing Unit

Base Clock 500 MHz 450 MHz
Boost Clock 1380 MHz 1296 MHz
Shading Units 5120 2048
Texture Mapping Units (TMUs) 320 128
Render Output Units (ROPs) 160 64
Compute Units (Pipelines) 640 256
Ray-tracing Cores - No
Instructions Per Cycle 2 IPC 2 IPC

Raw Performance

Pixel Fill Rate 221 GPixel/s 83 GPixel/s
Texture Fill Rate 442 GTexel/s 166 GTexel/s
FLOPS (FP32)
14.1 TFLOPS

Physical

Interface Custom Custom
TGP 60 W 30 W
Manufacturing TSMC TSMC
Fabrication Process 3 nm 5 nm
Transistor Count 56 billion 23.2 billion
Max. Temperature 100ยฐC 94ยฐC

Memory

Memory Type System Shared System Shared
Memory Clock 6400 MHz 6400 MHz
Effective Memory Speed 12800 Mbps 12800 Mbps
Bus 512-bit 256-bit
ECC No No
Memory Bandwidth
409.6 GB/s

API

Ray Tracing Yes No
DLSS No No

Cast your vote

Choose between two graphics cards
0 (0%)
0 (0%)
Total votes: < 1

User opinions

You can share your opinion or ask a question in the comments below
๐ŸŒ Register your profile and become part of NanoReview community!