Apple M4 Max GPU (32-core) vs GPU (10-Core)

We compared two integrated laptop professional GPUs: the Apple M4 Max GPU (32-core) with 512 pipelines and 4096 shaders against the 6 months older GPU (10-Core) that utilizes 160 pipelines and 1280 shaders. Here you will find complete details about specs, efficiency, performance tests, and more.

Review

General comparison of performance in games, applications, power efficiency, and other metrics
Gaming
Performance in DirectX, OpenCL, and Vulkan games
Workstation
Perf. in 3D modeling, video editing and rendering apps
AI/ML
Capabilities for machine learning and AI-related tasks
Energy Efficiency
Power consumption efficiency in different scenarios
NanoReview Final Score
Overall video card score

Key differences

Key distinctions and advantages of M4 GPU (10-Core) over M4 Max GPU (32-core)
Reasons to consider the Apple M4 Max GPU (32-core)
  • Performs significantly better (up to 3x) in 3DMark Steel Nomad Lite
  • 3.4x higher maximum theoretical performance (12.9 vs 3.8 TFLOPS)
  • Has 3.4x higher memory bandwidth: 409.6 vs 120 GB/s
  • Achieves 2.7x more points in the GeekBench 6 Compute test (100K vs 37K)
  • Has 3.2x more shading units (4096 vs 1280)

Benchmarks

Graphics cards’ performance in recent benchmarking apps

3D Mark

Multiplatform graphics benchmark suite that directly correlates with performance in modern games
Steel Nomad Lite Score
Solar Bay 50957 16554
Wild Life Extreme 30044 9610
Sources: 3DMark [1], [2]

GeekBench 6 OpenCL

GPU test for computational tasks (image processing, photography, computer vision, and ML)
GB6 Compute Score
Background Blur 174.8 img/sec 72.2 img/sec
Face Detection 112 img/sec 48.3 img/sec
Horizon Detection 4.19 Gpixels/sec 1.4 Gpixels/sec
Edge Detection 6.3 Gpixels/sec 1.94 Gpixels/sec
Gaussian Blur 4.69 Gpixels/sec 1.5 Gpixels/sec
Feature Matching 0.98 Gpixels/sec 0.53 Gpixels/sec
Stereo Matching 352.1 Gpixels/sec 123.3 Gpixels/sec
Particle Physics 14254.6 FPS 4938.4 FPS
API OpenCL OpenCL
Sources: Geekbench [3], [4]

Cinebench 2024 GPU

Hardware benchmark using Maxon's Cinema 4D rendering engine
Cinebench 2024 GPU

Blender

Rendering performance test for 3D modeling
Blender GPU
4437.57
Sources: Blender [9], [10]183 & 529 samples

Artificial Intelligence Tests

Performance in machine learning and artificial intelligence tasks

GeekBench 6 ML

Tests throughput of AI operations in single, half, and quantized precision
GB6 ML Single Precision
GB6 ML Half Precision
GB6 ML Quantized
Image Classification (SP) 8624 4874
Image Segmentation (HP) 20392 8168
Image Super Resolution (Q) 21752 11983
Face Detection (HP) 34099 17036
Pose Estimation (Q) 77846 31171
Text Classification (SP) 2993 2964
Machine Translation (HP) 5658 3410
Object Detection (SP) 7933 5112
Depth Estimation (Q) 37622 22700
Style Transfer (SP) 173882 60872
Framework Core ML Core ML
Backend GPU GPU
Sources: Geekbench [9], [10]

Specifications

Technical specifications of Apple M4 Max GPU (32-core) and GPU (10-Core)

General

Vendor Apple Apple
Build Integrated Integrated
Released October 30, 2024 May 7, 2024
Case Laptop Laptop
Purpose Professional Professional
Segment Mid-range Mid-range
Architecture Apple M GPU Apple M GPU
GPU Codename Custom Custom
Rival Equivalent - GeForce RTX 4070 Laptop - Adreno X1-85
Successor - Apple M5 Max GPU (32-core) - Apple M5 GPU (10-Core)
Recommended CPU - Apple M4 Max (14-Core) or above - Apple M4 (10-Core) or above
Used in CPUs - Apple M4 Max (14-Core) - Apple M4 (10-Core)
Laptop GPU ranking (18th and 71st place)

Graphics Processing Unit

Base Clock 500 MHz 500 MHz
Boost Clock 1578 MHz 1470 MHz
Shading Units 4096 1280
Texture Mapping Units (TMUs) 256 80
Render Output Units (ROPs) 128 40
Compute Units (Pipelines) 512 160
Instructions Per Cycle 2 IPC 2 IPC

Raw Performance

Pixel Fill Rate 202 GPixel/s 59 GPixel/s
Texture Fill Rate 404 GTexel/s 118 GTexel/s
FLOPS (FP32)
12.9 TFLOPS
3.8 TFLOPS

Physical

Interface Custom Custom
TGP 51 W 18 W
Manufacturing TSMC TSMC
Fabrication Process 3 nm 3 nm
Transistor Count - 28 billion
Max. Temperature 100°C 100°C

Memory

Memory Type System Shared System Shared
Memory Clock 8533 MHz 7500 MHz
Effective Memory Speed - 15000 Mbps
Bus 384-bit 128-bit
ECC No No
Memory Bandwidth
409.6 GB/s
120 GB/s

API

Ray Tracing Yes Yes
DLSS No No

Cast your vote

Choose between two graphics cards
3 (100%)
0 (0%)
Total votes: 3

User opinions

You can share your opinion or ask a question in the comments below
🌐 Register your profile and become part of NanoReview community!