PUNPCKLBW

/ VPUNPCKLBW

Interleave low bytes

Interleave the lower eight bytes of each 128-bit lane from two inputs.

SSE28-bit bits elements

Similar operations on Arm64:ZIP1vzip1q_u8

Instruction forms

3 in this sample
FormBitsRequiresEncoding
PUNPCKLBW xmm1, xmm2128SSE266 0F 60 /r
VPUNPCKLBW xmm1, xmm2, xmm3128AVXVEX.L0.66.0F.WIG 60 /r
VPUNPCKLBW ymm1, ymm2, ymm3256AVX2VEX.L1.66.0F.WIG 60 /r
128 bits · 16 × 8-bit elements

Explore the operation

The interactive explorer is loading. The operation and reference details are available below.

Operation

Pseudo-C · selected form
for each 128-bit lane:
  dst[2*i] = a[i]
  dst[2*i+1] = b[i]  // i = 0…7

Simplified pseudocode for the selected form.

Exact upstream semantics

Standard variant

let elements := 64 / element_size;
var result := Zero(64);
for i := 0 to elements-1 do
let j := i / 2;
if Is_Even(i) then
result[i *: element_size] := src1[j *: element_size];
else
result[i *: element_size] := src2[j *: element_size];
endif;
endfor;

Standard variant

let elements := 64 / element_size;
var result := Zero(64);
for i := 0 to elements-1 do
let j := i / 2;
if Is_Even(i) then
result[i *: element_size] := src1[j *: element_size];
else
result[i *: element_size] := src2[j *: element_size];
endif;
endfor;

Lane variant

let lanes := register_size / 128;
let elements := 128 / element_size;
var result := Zero(register_size);
for lane := 0 to lanes-1 do
let s1 := src1[lane *: 128];
let s2 := src2[lane *: 128];
var lane_result := Zero(128);
for i := 0 to elements-1 do
let j := i / 2;
let r := if Is_Even(i) then s1[j *: element_size] else s2[j *: element_size];
lane_result[i *: element_size] := r;
endfor;
result[lane *: 128] := lane_result;
endfor;

Parameterized by register and element size; from the pinned upstream definition.

What to watch for

  • The upper half of each input lane is unused.
  • The 256-bit form interleaves within each 128-bit lane.

Corresponding intrinsics

C / C++ · selected form
_mm_unpacklo_epi8
__m128i _mm_unpacklo_epi8(__m128i a, __m128i b)
#include <emmintrin.h>SSE2

Documented instruction mapping; a compiler may use an equivalent encoding or optimize the operation away.

Documented mapping

PUNPCKLBW xmm, xmm

Architectural details

Flags
Arithmetic flags unchanged.
Destination
Writes the low 128 bits; upper vector-register bits are preserved.
Encoding
66 0F 60 /r
Operands
  • xmm1read / write
  • xmm2read
Coverage
Selected register-only 128/256-bit forms. Memory operands, MMX, and EVEX/AVX-512 forms are outside this sample.
Exception information & execution requirements
  • #NM
  • #UD

Feature availability alone does not guarantee execution: OS state and execution-level controls also apply. Use the linked architecture documentation for the full exception conditions.

Performance measurements are not included in this preview. Latency and throughput depend on the exact form and microarchitecture. Measured data on uops.info ↗

Sources & provenance

Technical fields are imported from pinned upstream files. Explanations and explorer behavior are maintained separately.

Intel · PUNPCKLBW-PUNPCKLWD-PUNPCKLDQ-PUNPCKLQDQ-Unpack-Low-Data.xml
Revision
4ebe7f0ac1bd00f46244c49bb72c503ce368def7
SHA-256
2ff28acc17f0735fc26077d2d713c94fb7f6f909761d2b41a1f294239b228f67
Terms
Intel SDM terms; see License.md
View pinned upstream file ↗
Intel · intrinsics.xml
Revision
4ebe7f0ac1bd00f46244c49bb72c503ce368def7
SHA-256
6762a50652a35ef663fd532ed83b789d19bbc2ac7e94b1771f19d666d96c2c30
Terms
Intel SDM terms; see License.md
View pinned upstream file ↗

Intel sources are a documentation preview. Arm ACLE mappings are adapted under CC BY-SA 4.0 with an additional patent license. Arm MRS encoding data is distributed under BSD-3-Clause. © Arm Limited and contributors. Coverage and attribution.