Hand-coding a Linux ELF in raw machine code

Core requirement: Author the ELF headers and target-machine instruction bytes directly. This guide provides binary layouts, instruction encodings, and verification requirements. Deliver the executable, a portable runner, and behavioral tests.

Organizational adoption: the policy memo handles the paperwork. This guide handles the bytes.

Direct instruction encoding describes how the program is authored, not whether construction scripts exist. The supplied encoding references are part of the guide; construction utilities are not a required deliverable, and existing ones should be retained rather than hidden or deleted.

Minimal static ELF64

A static ET_EXEC ELF making raw syscalls needs a compatible Linux kernel and ISA (or emulation), executable permissions, and an execution policy that permits it. It needs no dynamic linker. The whole job is two structures then code. Non-PIE ET_EXEC loaded at a fixed 0x400000, so absolute addresses are known at build time.

ELF64 header (64 bytes):

fieldbytesvalue
e_ident16magic 7F 45 4C 46 (4 bytes) + 02 class=64-bit + 01 data=LE + 01 version + 00 OSABI=SysV + 00 ABIversion + 7 zero pad = 16. (Count: 4+1+1+1+1+1+7.)
e_type20200 = ET_EXEC
e_machine2x86-64 = 3E00 (0x3E); aarch64 = B700 (0xB7)
e_version401000000
e_entry80x400078 = load addr + sizeof(headers) = 0x400000 + 0x40 + 0x38
e_phoff80x40 (64)
e_shoff80 (no sections)
e_flags40
e_ehsize24000 (64)
e_phentsize23800 (56)
e_phnum20100 (1)
e_shentsize/e_shnum/e_shstrndx2+2+20,0,0

Program header `PT_LOAD` (56 bytes, ELF64 field order): p_type=1, p_flags=5 (R+X — note flags come second in ELF64), p_offset=0, p_vaddr=0x400000, p_paddr=0x400000, p_filesz=<file size>, p_memsz=<file size>, p_align=0x1000. Map the whole file from offset 0. Code starts at file offset 0x78 (= entry).

p_offset and p_vaddr must be congruent mod p_align — both are 0 mod 0x1000 here, so fine.

The two syscall ABIs (get these exactly right)

Syscall numbers differ between architectures. Check them alongside opcodes, operand fields, flags, and layout.

x86-64arm64 (aarch64)
syscall nr inraxx8
args inrdi, rsi, rdx, r10, r8, r9x0, x1, x2, x3, x4, x5
trap instructionsyscall = 0F 05svc #0 = D4000001
write164
exit6093
return valueraxx0

x86-64 instructions you need

mov r32, imm32 for the eight registers below is B8+r then a 4-byte LE immediate (unsigned bit pattern 0…0xFFFFFFFF; writing r32 zero-extends into r64). Extended registers require an additional REX prefix. Register codes for this form are 0…7 (eax=0,ecx=1,edx=2,ebx=3,esp=4,ebp=5,esi=6,edi=7):

B8 01000000   mov eax,1            ; write
BF 01000000   mov edi,1            ; fd=stdout
BE <addr32>   mov esi,&msg         ; abs addr works because ET_EXEC is non-PIE
BA 0E000000   mov edx,14           ; len
0F 05         syscall
B8 3C000000   mov eax,60           ; exit
BF 00000000   mov edi,0            ; status
0F 05         syscall

The string address is absolute: 0x400000 + file_offset_of_msg, written little-endian into the BE immediate. Note the string's offset is constant as long as the code in front of it is fixed-length — changing the message's text or length does NOT move the string (it always starts right after the same block of code), so the BE immediate stays the same; only edx (the length) changes. This example relies on the address fitting in 32 bits.

arm64 instructions you need

ARM64 instructions are fixed 32-bit, written little-endian (encoding 0xAABBCCDD → bytes DD CC BB AA). Encoding formulas (Rd = register number 0–31, hw = which 16-bit slot):

opformula (sum the disjoint pieces)
movz Xd,#imm16,lsl#(16*hw)0xD2800000 + (hw<<21) + (imm16<<5) + Rd
movk Xd,#imm16,lsl#(16*hw)0xF2800000 + (hw<<21) + (imm16<<5) + Rd
adr Xd,label0x10000000 + (immlo<<29) + (immhi<<5) + Rd, with off = label - addr(adr), immlo = off & 3, immhi = (off>>2) & 0x7FFFF
svc #imm160xD4000001 + (imm16<<5)

adr x1,msg is PC-relative (offset from the adr instruction to the string), so no absolute address is needed and it survives any load address. Syscall numbers ≤ 0xFFFF (like 64/93) load with a single movz; larger values need a movz for the low half then movk for the next.

D2800020   movz x0,#1           ; fd=stdout
100000E1   adr  x1,msg          ; &msg (PC-relative, off 0x1C here)
D28001C2   movz x2,#14          ; len
D2800808   movz x8,#64          ; write
D4000001   svc  #0
D2800000   movz x0,#0           ; status
D2800BA8   movz x8,#93          ; exit
D4000001   svc  #0

(movz x8,#64 = 0xD2800000 | (64<<5) | 8; movz x8,#93 likewise.)

Hello-world artifacts

These examples contain the literal header and instruction bytes. Their xxd -r -p streams contain only hex and whitespace: stray hex letters in comments would become bytes too. In these layouts the headers, code, and string are contiguous:

chmod +x the output. Inspect with readelf -h / readelf -l (or file).

Verify with qemu

qemu-user runs a foreign-architecture Linux binary on Linux by emulating the CPU and translating Linux syscalls. It is not an isolation boundary: run generated binaries only in an isolated Linux VM or container, including native-architecture tests.

On a Linux host: install qemu-user-static, then just:

qemu-x86_64 ./hello_x64      # or qemu-x86_64-static
qemu-aarch64 ./hello_arm64

A native-arch binary runs directly; only the foreign one needs the explicit qemu prefix (binfmt_misc usually makes even that automatic).

On macOS: Linux ELFs cannot run directly. verify-with-qemu.sh runs the hello examples with qemu-user inside a Docker Linux container. macOS does not ship QEMU; the runner image installs it inside Linux.

Always ship a portable runner (host arch ≠ target arch)

The host you build on is frequently a different arch than the binary (e.g. authoring an x86-64 ELF on an arm64 Mac). A bare ./bin then fails confusingly, and ad-hoc docker run … apt-get install qemu-user-static … one-liners reinstall qemu on every run and spew debconf/platform noise. So always drop a `run-elf.sh` next to the binary — a single entry point for supported Linux and macOS hosts with the prerequisites below. It does not provision application data, publish server ports, or sandbox native Linux execution.

Copy run-elf.sh into the project (or generate an equivalent). It:

./run-elf.sh ls_x64            # works on Linux x86-64, Linux arm64, or macOS — same command
./run-elf.sh ls_x64 | cat -v   # make ANSI escapes visible as text

Delivery and verification report

Deliver the executable, portable runner, and behavioral tests. Give copy-pasteable commands using the actual filenames, expected output and exit status, prerequisites, and application data requirements. Run generated binaries only inside an isolated Linux environment. Distinguish executed checks from skipped checks and known limitations; do not deploy as part of verification.

For the hello examples, ./run-elf.sh hello_arm64 and ./run-elf.sh hello_x64 should each print Hello, world! and exit 0. For the server, the repository's apps/exe/test.sh provisions /srv and checks real HTTP responses in temporary containers. The generic runner alone does not set up that doc root.

Beyond hello-world: servers, files, and big programs

hello-world is write+exit. Real programs — a TCP server, a file reader, a renderer — need more syscalls, writable memory, control flow, and subroutines, but the method is unchanged: you still author every instruction word by hand. This section is the delta. Worked example throughout: a web server that reads a markdown file from disk and renders it to HTML per request, entirely in hand-coded arm64 — a 7,016-byte static ELF that speaks HTTP and does its own markdown parsing.

More syscalls (arm64 / x86-64 numbers)

A blocking TCP server plus a file read use these (arm64 / x86-64):

callarm64x86-64callarm64x86-64
socket19841accept20243
setsockopt20854read630
bind20049close573
listen20150openat56257

Server shape: socket → setsockopt(SO_REUSEADDR) → bind → listen → loop{ accept → read(receive request bytes) → …work… → write → close }. sockaddr_in is 16 bytes: sin_family=2 (u16 LE), sin_port big-endian (port 8006 = 0x1F46, stored as bytes 1F 46; a little-endian halfword store would use the integer 0x461F to produce those bytes), sin_addr=0 (INADDR_ANY), 8 pad. TCP is a byte stream: one read does not guarantee a complete request, and writes may be short. Read a file with openat(AT_FDCWD, path, O_RDONLY) — AT_FDCWD is -100; load it with movn (arm64 movn x0,#99 = ~99 = -100 = 0x92800C60), path pointer in x1, flags 0 in x2. The path resolves against the process's cwd, so use an absolute path (/data.md) when the file's location is fixed — a bare data.md depends on where the container/process was started.

Writable memory: read() needs it, and BSS is free

read/accept write into a buffer, so that buffer must be in a writable page. Passing a buffer in an R+X-only PT_LOAD (p_flags=5) to read returns -EFAULT; a user-space store to that page instead faults. Buffers need writable storage, such as a separate RW segment or the stack. The existing server uses a single RWX segment (p_flags=7), which leaves code writable too. And big buffers cost zero file bytes: set `p_memsz` larger than `p_filesz` and the ELF loader zero-fills the memory beyond the file-backed bytes (classic BSS). The current server reserves a 4 MiB input buffer and a 256 KiB output buffer this way; that does not establish safe bounds for rendering arbitrary input. Buffers live past p_filesz but within p_memsz, at vaddrs you compute.

Guard syscall returns

A failed syscall returns a small negative number. Store read's return as an unsigned length without checking and -1 becomes a giant count — your copy/scan loop runs off the buffer and faults. Compare the return to 0 and branch negatives to an error path (e.g. emit a fixed HTTP/1.1 500), exactly like the fd check after openat.

More ARM64 instruction forms

bl sets the return address in x30; ret below returns through x30. Nested calls must preserve the return address; AAPCS64 callee-saved registers are x19–x29, and SP must remain 16-byte aligned at call boundaries and when used for memory access.

The following references cover only the exact forms shown: 64-bit arithmetic unless marked W, unshifted ADD/SUB immediates, unshifted register arithmetic, and unsigned scaled load/store offsets. Rd/Rn/Rm/Rt/Ra are five-bit fields, 0–31. Register 31 means SP or the zero register depending on the instruction and operand position, not an ordinary general-purpose register. The disjoint fields below are added with + (equivalently OR-ed after validation).

opencoding (sum the pieces)
mov Xd,Xm (alias of orr)0xAA0003E0 + (Rm<<16) + Rd
add Xd,Xn,#imm12 / sub0x91000000 / 0xD1000000, + (imm12<<10) + (Rn<<5) + Rd
cmp Xn,#imm12 (SUBS→xzr)0xF100001F + (imm12<<10) + (Rn<<5) (32-bit cmp Wn: 0x7100001F)
cmp Xn,Xm (SUBS→xzr)0xEB00001F + (Rm<<16) + (Rn<<5)
b.cond label0x54000000 + (imm19<<5) + cond; cond: eq0 ne1 cs2 cc3 mi4 pl5 hi8 ls9 ge10 lt11 gt12 le13
cbz Xt,label / cbnz0xB4000000 / 0xB5000000, + (imm19<<5) + Rt (32-bit: 0x34… / 0x35…)
b label / bl label0x14000000 / 0x94000000, + (imm26 & 0x3FFFFFF), imm26 = (label−addr(instruction))/4
ret0xD65F03C0
movn Xd,#imm160x92800000 + (imm16<<5) + Rd (#99 → −100 = AT_FDCWD)
ldrb Wt,[Xn,#imm] / strb0x39400000 / 0x39000000, + (imm<<10) + (Rn<<5) + Rt
strh Wt,[Xn,#imm]0x79000000 + ((imm>>1)<<10) + (Rn<<5) + Rt
ldr Wt,[Xn,#imm] / str0xB9400000 / 0xB9000000, + ((imm>>2)<<10) + (Rn<<5) + Rt
ldr Xt,[Xn,#imm] / str0xF9400000 / 0xF9000000, + ((imm>>3)<<10) + (Rn<<5) + Rt
udiv Xd,Xn,Xm0x9AC00800 + (Rm<<16) + (Rn<<5) + Rd
msub Xd,Xn,Xm,Xa0x9B008000 + (Rm<<16) + (Ra<<10) + (Rn<<5) + Rd (x−udiv*x → remainder)

Encoding and final-layout checks

Validate operands, register fields, immediate ranges, and required alignment before encoding. Masking is not validation. Check fixed opcode bits and operand-field placement too: address calculations are not the only source of errors.

Verify behavior, including the awkward inputs

Byte-for-byte comparison establishes agreement on the tested inputs, not universal correctness. A reference that shares the same algorithm, constants, or mistakes can agree perfectly and still be wrong. Examine its provenance and shared assumptions; supplement comparisons with independently specified expected outputs, boundary cases, malformed inputs, and structural properties such as Content-Length matching body size. Retain the tests and any reference needed to reproduce a claimed check.

The pilot's locally retained Python reference and construction script deliberately implement the same renderer algorithm. They are useful differential checks, not an independent proof of Markdown correctness. The repository tests include small literal expectations as a separate check, and report reference comparisons separately.

Running on macOS + Apple Silicon

An arm64 Linux ELF runs natively in a linux/arm64 Docker container on Apple Silicon — no qemu emulation, so it exercises the same ISA without CPU emulation; kernel versions, filesystem contents, resource limits, and security policy can still differ (a static ELF runs even in a busybox/scratch-style image; it needs no libc or dynamic linker). Two gotchas: (1) if docker pull/build hangs at error getting credentials, the credential helper may be stalling — bypass it for public images with an empty config: mkdir -p /tmp/nc && printf '{}' >/tmp/nc/config.json && DOCKER_CONFIG=/tmp/nc docker …. (2) A server never exits, so run it detached (docker run -d …) — a foreground docker run blocks your shell until killed.

Common mistakes