Apple M4 Max GPU (40-core) vs M1 Pro GPU (16-core)
We compared two integrated laptop professional GPUs: the Apple M4 Max GPU (40-core) with 640 pipelines and 5120 shaders against the 3 years and 1 month older M1 Pro GPU (16-core) that utilizes 256 pipelines and 2048 shaders. Here you will find complete details about specs, efficiency, performance tests, and more.
Review
Gaming
Performance in DirectX, OpenCL, and Vulkan games
Workstation
Perf. in 3D modeling, video editing and rendering apps
AI/ML
Capabilities for machine learning and AI-related tasks
Energy Efficiency
Power consumption efficiency in different scenarios
NanoReview Final Score
Overall video card score
Key differences
Reasons to consider the Apple M4 Max GPU (40-core)
- Performs significantly better (up to 3.6x) in 3DMark Steel Nomad Lite
- 3.1x higher maximum theoretical performance (16.2 vs 5.3 TFLOPS)
- Manufactured using a more efficient 3 nm process technology
- Achieves 2.8x more points in the GeekBench 6 Compute test (116K vs 42K)
- Has 2.7x higher memory bandwidth: 546 vs 204.8 GB/s
- Has 2.5x more shading units (5120 vs 2048)
Benchmarks
Graphics cardsโ performance in recent benchmarking apps3D Mark
Steel Nomad Lite Score
| Solar Bay | 61312 | 12485 |
| Wild Life Extreme | 36641 | 9873 |
GeekBench 6 OpenCL
GB6 Compute Score
| Background Blur | 195 img/sec | 74.5 img/sec |
| Face Detection | 125.9 img/sec | 47.3 img/sec |
| Horizon Detection | 4.82 Gpixels/sec | 1.73 Gpixels/sec |
| Edge Detection | 7.47 Gpixels/sec | 3.32 Gpixels/sec |
| Gaussian Blur | 5.8 Gpixels/sec | 1.75 Gpixels/sec |
| Feature Matching | 1.06 Gpixels/sec | 0.47 Gpixels/sec |
| Stereo Matching | 417.4 Gpixels/sec | 131.2 Gpixels/sec |
| Particle Physics | 17079.9 FPS | 5143 FPS |
| API | OpenCL | OpenCL |
Cinebench 2024 GPU
Cinebench 2024 GPU
Blender
Blender GPU
Recent User Tests
Apple M4 Max GPU (40-core)
| Date | Benchmark | Result |
|---|---|---|
| ๐ 2025-07-10 (thomas) | Geekbench 6 OpenCL | 116478 |
Apple M1 Pro GPU (16-core)
| Date | Benchmark | Result |
|---|---|---|
| ๐ 2026-03-27 (bbffx) | Geekbench 6 OpenCL | 43467 |
| ๐ 2026-03-27 (bbffx) | Cinebench 2024 | 2423 |
| ๐ 2026-03-27 (bbffx) | Steel Nomad Light | 4048 |
| ๐ 2025-02-21 (serhii) | Geekbench 6 OpenCL | 40932 |
Artificial Intelligence Tests
Performance in machine learning and artificial intelligence tasksGeekBench 6 ML
GB6 ML Single Precision
GB6 ML Half Precision
GB6 ML Quantized
| Image Classification (SP) | 10302 | 4112 |
| Image Segmentation (HP) | 23763 | 7488 |
| Image Super Resolution (Q) | 29314 | 9882 |
| Face Detection (HP) | 42016 | 14702 |
| Pose Estimation (Q) | 105972 | 33560 |
| Text Classification (SP) | 2963 | 2289 |
| Machine Translation (HP) | 6594 | 1629 |
| Object Detection (SP) | 8935 | 3633 |
| Depth Estimation (Q) | 43143 | 16986 |
| Style Transfer (SP) | 221563 | 53586 |
| Framework | Core ML | Core ML |
| Backend | GPU | GPU |
Specifications
Technical specifications of Apple M4 Max GPU (40-core) and M1 Pro GPU (16-core)General
| Vendor | Apple | Apple |
| Build | Integrated | Integrated |
| Released | October 30, 2024 | October 18, 2021 |
| Case | Laptop | Laptop |
| Purpose | Professional | Professional |
| Segment | Mid-range | Mid-range |
| Architecture | Apple M GPU | Apple M GPU |
| GPU Codename | Custom | - |
| Rival Equivalent | - GeForce RTX 4070 Laptop | - GeForce RTX 3050 Laptop |
| Successor | - Apple M5 Max GPU (40-core) | - Apple M5 Pro GPU (20-core) |
| Recommended CPU | - Apple M4 Max (16-Core) or above | - Apple M1 Pro or above |
| Used in CPUs | - Apple M4 Max (16-Core) | - Apple M1 Pro |
Laptop GPU ranking (14th and 77th place)
Graphics Processing Unit
| Base Clock | 500 MHz | 450 MHz |
| Boost Clock | 1578 MHz | 1296 MHz |
| Shading Units | 5120 | 2048 |
| Texture Mapping Units (TMUs) | 320 | 128 |
| Render Output Units (ROPs) | 160 | 64 |
| Compute Units (Pipelines) | 640 | 256 |
| Ray-tracing Cores | - | No |
| Instructions Per Cycle | 2 IPC | 2 IPC |
Raw Performance
| Pixel Fill Rate | 252 GPixel/s | 83 GPixel/s |
| Texture Fill Rate | 505 GTexel/s | 166 GTexel/s |
FLOPS (FP32)
Physical
| Interface | Custom | Custom |
| TGP | 50 W | 30 W |
| Manufacturing | TSMC | TSMC |
| Fabrication Process | 3 nm | 5 nm |
| Transistor Count | - | 23.2 billion |
| Max. Temperature | 100ยฐC | 94ยฐC |
Memory
| Memory Type | System Shared | System Shared |
| Memory Clock | 8533 MHz | 6400 MHz |
| Effective Memory Speed | - | 12800 Mbps |
| Bus | 512-bit | 256-bit |
| ECC | No | No |
Memory Bandwidth
API
| Ray Tracing | Yes | No |
| DLSS | No | No |
Cast your vote
Total votes: 7