NEON vs SVE: Fixed 128-Bit SIMD vs Scalable Vectors
Arm NEON and SVE both provide data-parallel execution, but NEON exposes fixed 128-bit vectors while SVE is built for scalable, predicated vector-length-agnostic code.
Arm NEON and SVE both accelerate data-parallel workloads, but they represent different generations of vector programming philosophy.
NEON exposes a fixed 128-bit SIMD model. SVE exposes scalable vectors whose implementation length is intentionally not hard-coded into portable software.
NEON’s fixed-width model
NEON vector registers are 128 bits wide in AArch64’s Advanced SIMD programming model. That means four 32-bit values or two 64-bit values fit in one 128-bit vector, for example.
Fixed width makes many intrinsics and data rearrangements straightforward because the lane count is known at compile time.
SVE’s scalable model
SVE allows hardware implementations to choose vector length within the architectural rules. Portable SVE loops use predicates and scalable progress rather than assuming a fixed number of elements per iteration.
This changes low-level loop structure substantially.
Tail handling
A NEON loop processing four 32-bit elements at a time may need scalar cleanup or another strategy when n is not divisible by four.
An SVE loop can generate a predicate for the final partial group and keep using the same predicated vector body.
Predication
SVE’s dedicated predicate registers are fundamental to the architecture. They allow element-wise activity control across loads, arithmetic and stores.
NEON provides conditional and selection techniques, but it was not designed around the same pervasive predicate-register model.
Does SVE replace NEON?
No. They can coexist in the Arm ecosystem, and support depends on the processor. NEON remains widely deployed and is often the appropriate SIMD target for software that needs compatibility across broad AArch64 hardware.
SVE is particularly attractive where scalable vector processing and its richer semantics are available and useful.
SVE2 broadens the scalable model
SVE2 adds operations aimed at a wider range of integer, DSP and media-style workloads while retaining the scalable SVE model. Some algorithms traditionally associated with NEON may therefore map naturally to SVE2 on processors that support it.
Intrinsics portability
NEON intrinsics describe fixed-size vector types and operations. SVE ACLE uses scalable vector and predicate types with restrictions appropriate to sizeless or scalable objects.
Porting hand-written NEON intrinsics to SVE is therefore not generally a mechanical “replace 128 with larger width” exercise. The loop should often be restructured around predication and vector-length-agnostic progress.
Auto-vectorization can hide the distinction
Well-written scalar loops may compile to NEON or SVE depending on target options, cost models and legality. This is one reason high-level code plus compiler diagnostics is worth trying before maintaining separate intrinsic implementations.
Performance comparisons need real processors
SVE’s scalable vector length means two SVE CPUs can differ in architectural vector size and execution resources. Likewise, NEON throughput differs across cores.
“SVE is wider, therefore faster” is not a valid general benchmark conclusion. Memory bandwidth, execution throughput, clocking and algorithm structure determine actual performance.
When to think NEON
- Broad compatibility with AArch64 CPUs is required.
- The algorithm maps naturally to fixed 128-bit chunks.
- A mature NEON implementation already exists and benchmarks well.
When to think SVE
- The target hardware supports SVE.
- Vector-length-agnostic portability is valuable.
- Predication, gather/scatter or scalable loop handling helps the algorithm.
- The compiler can exploit SVE effectively for the workload.
Key takeaways
- NEON is fixed-width 128-bit SIMD; SVE is scalable.
- SVE predicates make tail and conditional processing part of the normal vector model.
- SVE is not simply “wider NEON.”
- SVE2 extends the scalable model to more general-purpose data-processing operations.
- Choose based on target hardware, portability and measured performance.
References and further reading
Continue learning
Understand the loop. Understand the machine.
QCEV Vector Inspector is designed to help developers reason about vectorization opportunities, dependencies and memory access in performance-critical loops.