Apple M2 Max GPU (38-core) vs M1 Ultra GPU (64-core)

We compared two integrated laptop professional GPUs: the Apple M2 Max GPU (38-core) with 608 pipelines and 4864 shaders against the 10 months older M1 Ultra GPU (64-core) that utilizes 1024 pipelines and 8192 shaders. Here you will find complete details about specs, efficiency, performance tests, and more.

Review

General comparison of performance in games, applications, power efficiency, and other metrics
Gaming
Performance in DirectX, OpenCL, and Vulkan games
Workstation
Perf. in 3D modeling, video editing and rendering apps
AI/ML
Capabilities for machine learning and AI-related tasks
Energy Efficiency
Power consumption efficiency in different scenarios
NanoReview Final Score
Overall video card score

Key differences

Key distinctions and advantages of M1 Ultra GPU (64-core) over M2 Max GPU (38-core)
Reasons to consider the Apple M1 Ultra GPU (64-core)
  • Performs slightly better (up to 14%) in 3DMark Steel Nomad Lite
  • 56% higher maximum theoretical performance (21.2 vs 13.6 TFLOPS)
  • Has 2x higher memory bandwidth: 819.2 vs 409.6 GB/s
  • Achieves 22% more points in the GeekBench 6 Compute test (106K vs 86K)
  • Has 68% more shading units (8192 vs 4864)

Benchmarks

Graphics cards’ performance in recent benchmarking apps

3D Mark

Multiplatform graphics benchmark suite that directly correlates with performance in modern games
Steel Nomad Lite Score
Solar Bay 30368 36712
Wild Life Extreme 24948 29343
Sources: 3DMark [1], [2]

GeekBench 6 OpenCL

GPU test for computational tasks (image processing, photography, computer vision, and ML)
GB6 Compute Score
Background Blur 149 img/sec 157.3 img/sec
Face Detection 98.1 img/sec 108.7 img/sec
Horizon Detection 3.8 Gpixels/sec 5.07 Gpixels/sec
Edge Detection 6.88 Gpixels/sec 10.7 Gpixels/sec
Gaussian Blur 4.44 Gpixels/sec 6.59 Gpixels/sec
Feature Matching 0.73 Gpixels/sec 0.73 Gpixels/sec
Stereo Matching 264.9 Gpixels/sec 293.4 Gpixels/sec
Particle Physics 11144.2 FPS 14154.4 FPS
API OpenCL OpenCL
Sources: Geekbench [3], [4]

Cinebench 2024 GPU

Hardware benchmark using Maxon's Cinema 4D rendering engine
Cinebench 2024 GPU

Blender

Rendering performance test for 3D modeling
Blender GPU
Sources: Blender [9], [10] – 299 & 68 samples

Artificial Intelligence Tests

Performance in machine learning and artificial intelligence tasks

GeekBench 6 ML

Tests throughput of AI operations in single, half, and quantized precision
GB6 ML Single Precision
GB6 ML Half Precision
GB6 ML Quantized
Image Classification (SP) 6657 4834
Image Segmentation (HP) 13854 14458
Image Super Resolution (Q) 18642 13757
Face Detection (HP) 27965 23618
Pose Estimation (Q) 83144 80481
Text Classification (SP) 2377 2265
Machine Translation (HP) 2052 1798
Object Detection (SP) 6495 4975
Depth Estimation (Q) 31105 25755
Style Transfer (SP) 135686 135609
Framework Core ML Core ML
Backend GPU GPU
Sources: Geekbench [9], [10]

Specifications

Technical specifications of Apple M2 Max GPU (38-core) and M1 Ultra GPU (64-core)

General

Vendor Apple Apple
Build Integrated Integrated
Released January 17, 2023 March 8, 2022
Case Laptop Laptop
Purpose Professional Professional
Segment High-end High-end
Architecture Apple M GPU Apple M GPU
GPU Codename - Custom
Rival Equivalent - GeForce RTX 3060 Laptop - GeForce RTX 4070 Laptop
Successor - Apple M5 Max GPU (40-core) -
Recommended CPU - - Apple M1 Ultra or above
Used in CPUs - Apple M2 Max - Apple M1 Ultra
Laptop GPU ranking (41st and 31st place)

Graphics Processing Unit

Base Clock 450 MHz 450 MHz
Boost Clock 1398 MHz 1296 MHz
Shading Units 4864 8192
Texture Mapping Units (TMUs) 304 512
Render Output Units (ROPs) 152 256
Compute Units (Pipelines) 608 1024
Instructions Per Cycle 2 IPC 2 IPC

Raw Performance

Pixel Fill Rate 212 GPixel/s 332 GPixel/s
Texture Fill Rate 425 GTexel/s 664 GTexel/s
FLOPS (FP32)
13.6 TFLOPS
21.2 TFLOPS

Physical

Interface Custom Custom
TGP 70 W 120 W
Manufacturing TSMC TSMC
Fabrication Process 5 nm 5 nm
Transistor Count 52 billion 92.8 billion
Max. Temperature 94°C 94°C

Memory

Memory Type System Shared System Shared
Memory Clock 6400 MHz 6400 MHz
Effective Memory Speed 12800 Mbps 12800 Mbps
Bus 512-bit 1024-bit
ECC No No
Memory Bandwidth
819.2 GB/s

API

Ray Tracing No No
DLSS No No

Cast your vote

Choose between two graphics cards
0 (0%)
4 (100%)
Total votes: 4

User opinions

You can share your opinion or ask a question in the comments below
🌐 Register your profile and become part of NanoReview community!