Hand-coding a Linux ELF in raw machine code
Core requirement: Author the ELF headers and target-machine instruction bytes directly. This guide provides binary layouts, instruction encodings, and verification requirements. Deliver the executable, a portable runner, and behavioral tests.
Organizational adoption: the policy memo handles the paperwork. This guide handles the bytes.
Direct instruction encoding describes how the program is authored, not whether construction scripts exist. The supplied encoding references are part of the guide; construction utilities are not a required deliverable, and existing ones should be retained rather than hidden or deleted.
Minimal static ELF64
A static ET_EXEC ELF making raw syscalls needs a compatible Linux kernel and ISA (or emulation), executable permissions, and an execution policy that permits it. It needs no dynamic linker. The whole job is two structures then code. Non-PIE ET_EXEC loaded at a fixed 0x400000, so absolute addresses are known at build time.
ELF64 header (64 bytes):
| field | bytes | value |
|---|---|---|
| e_ident | 16 | magic 7F 45 4C 46 (4 bytes) + 02 class=64-bit + 01 data=LE + 01 version + 00 OSABI=SysV + 00 ABIversion + 7 zero pad = 16. (Count: 4+1+1+1+1+1+7.) |
| e_type | 2 | 0200 = ET_EXEC |
| e_machine | 2 | x86-64 = 3E00 (0x3E); aarch64 = B700 (0xB7) |
| e_version | 4 | 01000000 |
| e_entry | 8 | 0x400078 = load addr + sizeof(headers) = 0x400000 + 0x40 + 0x38 |
| e_phoff | 8 | 0x40 (64) |
| e_shoff | 8 | 0 (no sections) |
| e_flags | 4 | 0 |
| e_ehsize | 2 | 4000 (64) |
| e_phentsize | 2 | 3800 (56) |
| e_phnum | 2 | 0100 (1) |
| e_shentsize/e_shnum/e_shstrndx | 2+2+2 | 0,0,0 |
Program header `PT_LOAD` (56 bytes, ELF64 field order): p_type=1, p_flags=5 (R+X — note flags come second in ELF64), p_offset=0, p_vaddr=0x400000, p_paddr=0x400000, p_filesz=<file size>, p_memsz=<file size>, p_align=0x1000. Map the whole file from offset 0. Code starts at file offset 0x78 (= entry).
p_offset and p_vaddr must be congruent mod p_align — both are 0 mod 0x1000 here, so fine.
The two syscall ABIs (get these exactly right)
Syscall numbers differ between architectures. Check them alongside opcodes, operand fields, flags, and layout.
| x86-64 | arm64 (aarch64) | |
|---|---|---|
| syscall nr in | rax | x8 |
| args in | rdi, rsi, rdx, r10, r8, r9 | x0, x1, x2, x3, x4, x5 |
| trap instruction | syscall = 0F 05 | svc #0 = D4000001 |
| write | 1 | 64 |
| exit | 60 | 93 |
| return value | rax | x0 |
x86-64 instructions you need
mov r32, imm32 for the eight registers below is B8+r then a 4-byte LE immediate (unsigned bit pattern 0…0xFFFFFFFF; writing r32 zero-extends into r64). Extended registers require an additional REX prefix. Register codes for this form are 0…7 (eax=0,ecx=1,edx=2,ebx=3,esp=4,ebp=5,esi=6,edi=7):
B8 01000000 mov eax,1 ; write
BF 01000000 mov edi,1 ; fd=stdout
BE <addr32> mov esi,&msg ; abs addr works because ET_EXEC is non-PIE
BA 0E000000 mov edx,14 ; len
0F 05 syscall
B8 3C000000 mov eax,60 ; exit
BF 00000000 mov edi,0 ; status
0F 05 syscall
The string address is absolute: 0x400000 + file_offset_of_msg, written little-endian into the BE immediate. Note the string's offset is constant as long as the code in front of it is fixed-length — changing the message's text or length does NOT move the string (it always starts right after the same block of code), so the BE immediate stays the same; only edx (the length) changes. This example relies on the address fitting in 32 bits.
arm64 instructions you need
ARM64 instructions are fixed 32-bit, written little-endian (encoding 0xAABBCCDD → bytes DD CC BB AA). Encoding formulas (Rd = register number 0–31, hw = which 16-bit slot):
| op | formula (sum the disjoint pieces) |
|---|---|
movz Xd,#imm16,lsl#(16*hw) | 0xD2800000 + (hw<<21) + (imm16<<5) + Rd |
movk Xd,#imm16,lsl#(16*hw) | 0xF2800000 + (hw<<21) + (imm16<<5) + Rd |
adr Xd,label | 0x10000000 + (immlo<<29) + (immhi<<5) + Rd, with off = label - addr(adr), immlo = off & 3, immhi = (off>>2) & 0x7FFFF |
svc #imm16 | 0xD4000001 + (imm16<<5) |
adr x1,msg is PC-relative (offset from the adr instruction to the string), so no absolute address is needed and it survives any load address. Syscall numbers ≤ 0xFFFF (like 64/93) load with a single movz; larger values need a movz for the low half then movk for the next.
D2800020 movz x0,#1 ; fd=stdout
100000E1 adr x1,msg ; &msg (PC-relative, off 0x1C here)
D28001C2 movz x2,#14 ; len
D2800808 movz x8,#64 ; write
D4000001 svc #0
D2800000 movz x0,#0 ; status
D2800BA8 movz x8,#93 ; exit
D4000001 svc #0
(movz x8,#64 = 0xD2800000 | (64<<5) | 8; movz x8,#93 likewise.)
Hello-world artifacts
These examples contain the literal header and instruction bytes. Their xxd -r -p streams contain only hex and whitespace: stray hex letters in comments would become bytes too. In these layouts the headers, code, and string are contiguous:
chmod +x the output. Inspect with readelf -h / readelf -l (or file).
Verify with qemu
qemu-user runs a foreign-architecture Linux binary on Linux by emulating the CPU and translating Linux syscalls. It is not an isolation boundary: run generated binaries only in an isolated Linux VM or container, including native-architecture tests.
On a Linux host: install qemu-user-static, then just:
qemu-x86_64 ./hello_x64 # or qemu-x86_64-static
qemu-aarch64 ./hello_arm64
A native-arch binary runs directly; only the foreign one needs the explicit qemu prefix (binfmt_misc usually makes even that automatic).
On macOS: Linux ELFs cannot run directly. verify-with-qemu.sh runs the hello examples with qemu-user inside a Docker Linux container. macOS does not ship QEMU; the runner image installs it inside Linux.
Always ship a portable runner (host arch ≠ target arch)
The host you build on is frequently a different arch than the binary (e.g. authoring an x86-64 ELF on an arm64 Mac). A bare ./bin then fails confusingly, and ad-hoc docker run … apt-get install qemu-user-static … one-liners reinstall qemu on every run and spew debconf/platform noise. So always drop a `run-elf.sh` next to the binary — a single entry point for supported Linux and macOS hosts with the prerequisites below. It does not provision application data, publish server ports, or sandbox native Linux execution.
Copy run-elf.sh into the project (or generate an equivalent). It:
- reads the target arch from the ELF's own two-byte
e_machinefield (offset0x12:3e00=x86-64,b700=arm64) — no hardcoding, - checks ELF64 little-endian identification and runs natively on Linux when host arch == target,
- uses `qemu-<arch>-static` or `qemu-<arch>` when Linux has a different architecture (requires qemu-user),
- on macOS builds a small cached Docker image with
qemu-user-staticonce, then reuses it and passes a TTY when stdin and stdout are terminals. Docker with a Linux VM is required.
./run-elf.sh ls_x64 # works on Linux x86-64, Linux arm64, or macOS — same command
./run-elf.sh ls_x64 | cat -v # make ANSI escapes visible as text
Delivery and verification report
Deliver the executable, portable runner, and behavioral tests. Give copy-pasteable commands using the actual filenames, expected output and exit status, prerequisites, and application data requirements. Run generated binaries only inside an isolated Linux environment. Distinguish executed checks from skipped checks and known limitations; do not deploy as part of verification.
For the hello examples, ./run-elf.sh hello_arm64 and ./run-elf.sh hello_x64 should each print Hello, world! and exit 0. For the server, the repository's apps/exe/test.sh provisions /srv and checks real HTTP responses in temporary containers. The generic runner alone does not set up that doc root.
Beyond hello-world: servers, files, and big programs
hello-world is write+exit. Real programs — a TCP server, a file reader, a renderer — need more syscalls, writable memory, control flow, and subroutines, but the method is unchanged: you still author every instruction word by hand. This section is the delta. Worked example throughout: a web server that reads a markdown file from disk and renders it to HTML per request, entirely in hand-coded arm64 — a 7,016-byte static ELF that speaks HTTP and does its own markdown parsing.
More syscalls (arm64 / x86-64 numbers)
A blocking TCP server plus a file read use these (arm64 / x86-64):
| call | arm64 | x86-64 | call | arm64 | x86-64 | |
|---|---|---|---|---|---|---|
| socket | 198 | 41 | accept | 202 | 43 | |
| setsockopt | 208 | 54 | read | 63 | 0 | |
| bind | 200 | 49 | close | 57 | 3 | |
| listen | 201 | 50 | openat | 56 | 257 |
Server shape: socket → setsockopt(SO_REUSEADDR) → bind → listen → loop{ accept → read(receive request bytes) → …work… → write → close }. sockaddr_in is 16 bytes: sin_family=2 (u16 LE), sin_port big-endian (port 8006 = 0x1F46, stored as bytes 1F 46; a little-endian halfword store would use the integer 0x461F to produce those bytes), sin_addr=0 (INADDR_ANY), 8 pad. TCP is a byte stream: one read does not guarantee a complete request, and writes may be short. Read a file with openat(AT_FDCWD, path, O_RDONLY) — AT_FDCWD is -100; load it with movn (arm64 movn x0,#99 = ~99 = -100 = 0x92800C60), path pointer in x1, flags 0 in x2. The path resolves against the process's cwd, so use an absolute path (/data.md) when the file's location is fixed — a bare data.md depends on where the container/process was started.
Writable memory: read() needs it, and BSS is free
read/accept write into a buffer, so that buffer must be in a writable page. Passing a buffer in an R+X-only PT_LOAD (p_flags=5) to read returns -EFAULT; a user-space store to that page instead faults. Buffers need writable storage, such as a separate RW segment or the stack. The existing server uses a single RWX segment (p_flags=7), which leaves code writable too. And big buffers cost zero file bytes: set `p_memsz` larger than `p_filesz` and the ELF loader zero-fills the memory beyond the file-backed bytes (classic BSS). The current server reserves a 4 MiB input buffer and a 256 KiB output buffer this way; that does not establish safe bounds for rendering arbitrary input. Buffers live past p_filesz but within p_memsz, at vaddrs you compute.
Guard syscall returns
A failed syscall returns a small negative number. Store read's return as an unsigned length without checking and -1 becomes a giant count — your copy/scan loop runs off the buffer and faults. Compare the return to 0 and branch negatives to an error path (e.g. emit a fixed HTTP/1.1 500), exactly like the fd check after openat.
More ARM64 instruction forms
bl sets the return address in x30; ret below returns through x30. Nested calls must preserve the return address; AAPCS64 callee-saved registers are x19–x29, and SP must remain 16-byte aligned at call boundaries and when used for memory access.
The following references cover only the exact forms shown: 64-bit arithmetic unless marked W, unshifted ADD/SUB immediates, unshifted register arithmetic, and unsigned scaled load/store offsets. Rd/Rn/Rm/Rt/Ra are five-bit fields, 0–31. Register 31 means SP or the zero register depending on the instruction and operand position, not an ordinary general-purpose register. The disjoint fields below are added with + (equivalently OR-ed after validation).
| op | encoding (sum the pieces) |
|---|---|
mov Xd,Xm (alias of orr) | 0xAA0003E0 + (Rm<<16) + Rd |
add Xd,Xn,#imm12 / sub | 0x91000000 / 0xD1000000, + (imm12<<10) + (Rn<<5) + Rd |
cmp Xn,#imm12 (SUBS→xzr) | 0xF100001F + (imm12<<10) + (Rn<<5) (32-bit cmp Wn: 0x7100001F) |
cmp Xn,Xm (SUBS→xzr) | 0xEB00001F + (Rm<<16) + (Rn<<5) |
b.cond label | 0x54000000 + (imm19<<5) + cond; cond: eq0 ne1 cs2 cc3 mi4 pl5 hi8 ls9 ge10 lt11 gt12 le13 |
cbz Xt,label / cbnz | 0xB4000000 / 0xB5000000, + (imm19<<5) + Rt (32-bit: 0x34… / 0x35…) |
b label / bl label | 0x14000000 / 0x94000000, + (imm26 & 0x3FFFFFF), imm26 = (label−addr(instruction))/4 |
ret | 0xD65F03C0 |
movn Xd,#imm16 | 0x92800000 + (imm16<<5) + Rd (#99 → −100 = AT_FDCWD) |
ldrb Wt,[Xn,#imm] / strb | 0x39400000 / 0x39000000, + (imm<<10) + (Rn<<5) + Rt |
strh Wt,[Xn,#imm] | 0x79000000 + ((imm>>1)<<10) + (Rn<<5) + Rt |
ldr Wt,[Xn,#imm] / str | 0xB9400000 / 0xB9000000, + ((imm>>2)<<10) + (Rn<<5) + Rt |
ldr Xt,[Xn,#imm] / str | 0xF9400000 / 0xF9000000, + ((imm>>3)<<10) + (Rn<<5) + Rt |
udiv Xd,Xn,Xm | 0x9AC00800 + (Rm<<16) + (Rn<<5) + Rd |
msub Xd,Xn,Xm,Xa | 0x9B008000 + (Rm<<16) + (Ra<<10) + (Rn<<5) + Rd (x−udiv*x → remainder) |
Encoding and final-layout checks
Validate operands, register fields, immediate ranges, and required alignment before encoding. Masking is not validation. Check fixed opcode bits and operand-field placement too: address calculations are not the only source of errors.
- For the X-register
movz/movkforms above:imm16is 0…65535 andhwis 0…3. The shownmovnhashw=0;svchas an unsigned 16-bit immediate. W-register wide moves allow onlyhw=0…1and use different opcode bits. - The shown ADD/SUB/CMP immediate forms have an unshifted unsigned 12-bit immediate (0…4095). This is not the format for shifted-immediate or register variants. The shown register CMP has shift amount zero.
- For
adr, subtract the address of the ADR instruction itself from the data address. Its signed 21-bit displacement is measured in bytes: −1048576…1048575, with no target-alignment restriction. Split into immlo/immhi only after checking that range. This is not ADRP's page-relative format. - For
b/bl, subtract the address of the branch instruction itself from its target, require divisibility by 4, and check the signed 26-bit scaled displacement: byte range −134217728…134217724. Forb.condand the shown W/Xcbz/cbnz, the signed field is 19 bits: byte range −1048576…1048572, also scaled by 4. Encode the checked signed field in two's complement. Validate the condition code forb.cond(0…15; the named conditions in the table are a subset). These rules do not describe TBZ/TBNZ. - For the unsigned-offset LDR/STR forms above, the field is unsigned 12-bit: byte offsets must be nonnegative, divisible by the access size, and at most 4095 × that size. Byte: 0…4095; halfword: 0…8190 by 2; W: 0…16380 by 4; X: 0…32760 by 8. Divide only after validation. Unscaled, pre/post-indexed, pair, and register-offset forms have different rules.
- The retained construction utility also uses LDURB (signed unscaled imm9, −256…255 bytes), X-register STP pre-index/LDP post-index (signed imm7 × 8, −512…504 bytes), and the X-register LSL alias (shift 0…63). For writeback pairs reject base/transfer-register overlap except SP, and for LDP reject identical destination registers. A representable SP adjustment must still preserve the required stack alignment.
- ARM64 instruction addresses and direct branch targets must be four-byte aligned. Data labels follow their own alignment requirements; strings and other byte data need not be four-byte aligned.
- Resolve every referenced address and displacement against the final layout. Reject unresolved or ambiguous references, including duplicate label definitions. If existing tooling uses multiple passes, compare label positions and instruction locations as well as sizes; equal total size alone cannot detect a moved label.
- Test representable limits, values just outside them, misaligned values, invalid registers, and known opcode/operand-field vectors. Inspect the final ELF's entry, segments, file/memory bounds, permissions, and alignment as well as the instructions.
Verify behavior, including the awkward inputs
Byte-for-byte comparison establishes agreement on the tested inputs, not universal correctness. A reference that shares the same algorithm, constants, or mistakes can agree perfectly and still be wrong. Examine its provenance and shared assumptions; supplement comparisons with independently specified expected outputs, boundary cases, malformed inputs, and structural properties such as Content-Length matching body size. Retain the tests and any reference needed to reproduce a claimed check.
The pilot's locally retained Python reference and construction script deliberately implement the same renderer algorithm. They are useful differential checks, not an independent proof of Markdown correctness. The repository tests include small literal expectations as a separate check, and report reference comparisons separately.
Running on macOS + Apple Silicon
An arm64 Linux ELF runs natively in a linux/arm64 Docker container on Apple Silicon — no qemu emulation, so it exercises the same ISA without CPU emulation; kernel versions, filesystem contents, resource limits, and security policy can still differ (a static ELF runs even in a busybox/scratch-style image; it needs no libc or dynamic linker). Two gotchas: (1) if docker pull/build hangs at error getting credentials, the credential helper may be stalling — bypass it for public images with an empty config: mkdir -p /tmp/nc && printf '{}' >/tmp/nc/config.json && DOCKER_CONFIG=/tmp/nc docker …. (2) A server never exits, so run it detached (docker run -d …) — a foreground docker run blocks your shell until killed.
Common mistakes
- Using x86-64 syscall numbers on arm64 or vice-versa (write is 1 vs 64, exit 60 vs 93) — the program does the wrong thing or hangs.
- Passing read-only memory as an output buffer to a syscall — the buffer must be writable.
- Using a syscall return as an unsigned length without checking for negative — one errno becomes a giant length and overruns the buffer.
- Shipping big zero buffers as file bytes — use
p_memsz > p_filesz(BSS zero-fill) instead of padding megabytes of zeros into the ELF. - Treating agreement with a shared-algorithm reference as universal correctness — test independent expectations and boundaries too.
syscall(0F 05) on x86-64 vssvc #0(D4000001) on arm64 — don't cross them. Linux arm64 issvc #0with the number inx8.- Wrong
e_machine(0x3E vs 0xB7) — the kernel/qemu rejects "Exec format error". - Hardcoding a
mov esi,imm32absolute string address but building PIE — keep itET_EXECso0x400000is fixed. - Forgetting ELF64 puts
p_flagsimmediately afterp_type(different from ELF32). - Trying
qemu-system-*for a static binary on macOS — useqemu-user-staticin a container instead.