| Age | Commit message (Collapse) | Author |
|
The resident firmware layer real machines ship, the thing the
gic group lesson pointed at. The reset path configures EL3,
SP_EL3 on its own region, the monitor vectors in VBAR_EL3,
then hands the next stage non-secure EL2 in the manual's boot
state and never comes back except through exceptions.
Secondaries that enter at EL3 get the monitor before they
park, a firmware call on any PE must land in a handler, and
SCR_EL3.NS is set to match the primary so a released PE does
not come up secure while the kernel runs non-secure.
The SMC conduit traps into the lower EL AArch64 sync slot and
dispatches through the same PSCI C code the hvc path uses,
SMCCC register convention kept whole across the trap.
On the emulator here the machine's own firmware shadow stands
in front of the conduit, its PSCI answers before the monitor
sees the call, and its secure memory map traps the kernel's
flash probe after init starts. The monitor mechanics, the
entry, the vectors, the stack, the eret, the SMC layout, are
live on every secure boot, the call dispatch itself is the
hardware receipt.
receipt: secure boot through the monitor to four cpus and the
init exec, plain boot unchanged to the busybox shell.
|
|
The interrupt controller state the kernel inherits is the
bootloader's to define. On real hardware the secure world
owns which interrupts the non-secure kernel will ever see,
and a distributor left with random enables or secure group
bits can fire before the kernel's irqchip driver is up. The
gic goes into the defined state here, distributor off, every
line in the non-secure group, per interrupt enables, pending
and active cleared, cpu interface off. The kernel programs
everything it runs with itself, it starts from zero instead
of from whatever the last stage left.
The controller is found in the devicetree, no hardcoded
address. The new walker locates a node by name at any depth
and decodes the first reg pair with the root cell counts,
the same parse the kernel does. The node's own begin token
starts the walk at depth zero, starting at one skips every
prop in the node, the first version matched nothing.
receipt: gic 8000000 off, root irq handler gic_handle_irq,
smp brought up 1 node 4 cpus, run /init, busybox shell, the
uart console driven by irq 14 through the gic the kernel
reprogrammed over our off state.
|
|
Firmware edits the devicetree it hands the kernel: new
properties cannot go into a packed fdt in place, the blob moves
to scratch first and the blocks grow there. The insert moves the
strings block up by the prop size, opens the prop slot before
/chosen's end token, the new name lands at the strings end, and
all four header fields track the geometry, totalsize, off_dt_
strings, size_dt_strings and size_dt_struct. The last one bounds
the token walk in libfdt and a stale value there reads as
BADSTRUCTURE, the kernel parsed no memory node and panicked on
page table allocation before the first print.
The strings block length comes from the size_dt_strings header
field, not totalsize minus off_dt_strings, qemu's tree carries
a hole after the strings block and the subtraction drags it
along as tree.
The initrd start and end land in /chosen as real properties now,
the initrd= bootargs shortcut is gone.
receipt: smp brought up 1 node 4 cpus, run /init, userspace,
busybox shell, the initrd mounted from the chosen cells.
|
|
qemu -kernel parses a raw arm64 blob as a linux Image and enters
at RAMBASE plus whatever text_offset it guesses out of the
garbage, 0x80000 in our case. every wild PC at image+0x80000 in
the debug logs was our own code running from the wrong address.
the binary now carries a real Image header: code0 branches over
it, magic ARM\x64 at 0x38, text_offset 0, image_size stamped
after objcopy by tools/fillsize.py.
the runtime also split by exception level. the C body runs at
EL1, the semihosting hlt is answered by qemu only from EL2, so
the EL2 vector replays the trap there and erets home with the
result. the kernel handoff hvc raises back to EL2 where
booting.rst wants it, the same vector slot dispatches PSCI hvc
from the kernel, boot handoff and semihosting by EC and function
id.
the payload load address was hardcoded 0x40200000, which is where
qemu placed our image, so the load overwrote the running
bootloader with kernel bytes mid flight. the load address is now
__image_copy_end plus 16MB, wherever the image actually runs.
receipt: run /init, tashaboot linux userspace reached, cores: 4,
busybox shell on a 4 cpu virt machine with initrd.
|
|
The bootloader now does the whole job of machine firmware it owns:
boots 4 cpus, hands over an initrd, answers PSCI, and carries the
delay and cache primitives the arch layer needs.
SMP: the secondary pen is the Wait For Event mechanism from the
manual (B2-144, D1-2255), each secondary watches its spin gate,
WFE, the release writes the entry and SEVs, the recheck after each
wake covers a release that lands between the load and the sleep.
The gates land in the dtb cpu-release-addr slots, rewritten in
place by a small walker, no libfdt, structure per the devicetree
specification, values only, the properties themselves are fixed at
build time like firmware shipping a fixed blob.
PSCI 0.2 at EL2 (DEN 0022): the HVC trap arrives at the current EL
SP_ELx sync slot (EC 0x16 in ESR_EL2, the vector layout Table D1-7),
dispatch on the standard function ids, VERSION, CPU_ON writes the
target gate and SEVs, CPU_OFF clears the gate and returns to the
pen, SYSTEM_OFF and SYSTEM_RESET drive RMR_EL2.RR. On qemu the cores
are held by the machine's own firmware and released through its PSCI
(hvc with -kernel, smc with virtualization=on), the handler here is
the real hardware path where the bootloader is the conduit.
The initrd handoff: loaded at a fixed address clear of the image
and dtb, the dtb /chosen carries linux,initrd-start and -end.
Delays are the generic timer (D10), CNTFRQ_EL0 frequency, CNTVCT_EL0
count, busy wait, no interrupts. Cache maintenance by virtual
address, dc cvac, dc ivac, dc civac, ic ivau with the barrier pairs
the manual requires, the by VA form beats set and way when the
range is known.
Boot receipt, 4 cpus, el2, initrd:
tashaboot 0.1
initrd at 46000000
[ 0.000000] Booting Linux on physical CPU 0x0000000000
[ 0.130621] smp: Brought up 1 node, 4 CPUs
[ 1.830692] Run /init as init process
tashaboot linux userspace reached
cores: 4
BusyBox v1.37.0 built-in shell (ash)
~ #
Signed-off-by: Bradley Morgan <brads@mainlining.org>
|
|
The bootloader now builds its own stage 1 translation tables instead
of only tearing firmware state down. one L0 table, one L1 under it,
device nGnRE block for the low 1GB, normal writeback 2MB blocks for
RAM. the descriptors, attribute encodings, MAIR and TCR settings come
straight from the manual, level 0/1/2 and level 3 formats at D5-2444
and D5-2447, stage 1 attribute fields at D5-2451, MAIR region
attributes at D5-2476, the PA size from ID_AA64MMFR0_EL1.PARange per
D5-2399.
The tables are EL aware, TTBR0/TCR/MAIR at whichever regime the entry
left us in, EL2 or EL1, and the self test translates through AT
S1E2R or AT S1E1R per the exception level and checks PAR_EL1 for the
identity result:
mmu: mmio 0x09000000 (uart) ok, pa 9000000
mmu: mmio 0x00000000 ok, pa 0
mmu: ram 0x40200000 (load) ok, pa 40200000
mmu: ram 0x41000000 ok, pa 41000000
mmu: self 0x40080000 ok, pa 40080000
mmu: identity map on
The map is torn down again before the payload, the kernel wants the
architecture state at entry, not ours.
Two bugs the self test caught on the way. T0SZ was 25 for a 39-bit
VA, but with the 4KB granule a 39-bit VA starts the walk at level 1,
and the L0 indexed structure was misread one level over, every
descriptor landed in the wrong slot and all fetches past the first
2MB faulted level 1. T0SZ is 16 now, the walk starts at level 0 and
the three level structure matches. The second, the mmio table was
orphaned, the l0 entry was written twice and the second write won,
so the device block was never reachable and AT on the uart address
faulted. the mmio block now lives at l1[0] in the same L1 table as
RAM.
The stack also moved to its own region above the bss in the linker
script. the tables are bss objects, a stack growing down from the
bss end shares their address space and a deep call chain writes into
the top table.
Signed-off-by: Bradley Morgan <brads@mainlining.org>
|
|
A small arm64 bootloader. No board code, no device tree porting, the
architecture manual is the whole story: exception vectors in the
fixed 16 slot layout (Table D1-7), ESR_ELx decoded by exception class
(D1-2172), EL entry and eret chains per the programmers model
(D1-2146), cache maintenance by set/way over the CLIDR_EL1 levels,
semihosting for console and file io per DUI 0203, and the A64 boot
protocol from Documentation/arch/arm64/booting.rst.
The loader boots a stock mainline Image end to end on the qemu virt
machine. Boot receipt with 7.3-rc3 (42MB Image):
tashaboot 0.1
loaded 43450368 bytes at 40200000, entry 40200000
jumping
[ 0.000000] Booting Linux on physical CPU 0x0000000000 [0x411fd070]
[ 0.000000] Linux version 7.3.0-rc3
[ 0.000000] Machine model: linux,dummy-virt
[ 0.000000] earlycon: pl11 MMIO32:0x0000000009000000
...
---[ end Kernel panic - not syncing: VFS: Unable to mount root fs ]---
The panic is the expected end state, no root filesystem is handed
over yet.
The boot chain, state per stage, start to payload:
+-----------+-----+--------------+----------------------------------+
| stage | EL | state | work |
+-----------+-----+--------------+----------------------------------+
| firmware | any | MMU maybe on | x0 = dtb, jump in |
+-----------+-----+--------------+----------------------------------+
| tashaboot | 3-2 | | SCR_EL3.NS = 1, eret to EL2 |
+-----------+-----+--------------+----------------------------------+
| | 2 | virt scrub | HCR/CNTHCTL/CPTR/HSTR, CNTFRQ, |
| | | | VBAR_EL2, MMU off, tlbi alle2 |
+-----------+-----+--------------+----------------------------------+
| | 2 | | load Image over semihosting, |
| | | | validate header, place per |
| | | | booting.rst |
+-----------+-----+--------------+----------------------------------+
| | 2 | caches clean | flush dcache, inval icache, |
| | | | args ride x20/x21, regs last |
+-----------+-----+--------------+----------------------------------+
| payload | 2 | fresh start | x0 = dtb, x1-x3 = 0, DAIF |
| | | | masked, br to image entry |
+-----------+-----+--------------+----------------------------------+
Two handoff bugs the kernel caught, both AAPCS clobbers in the final
jump. Cache maintenance was called after the register setup, x0-x18
are caller saved, so tb_flush_dcache_all() wiped the dtb pointer and
the kernel spun in setup_machine_fdt() with an invalid device tree
blob. The flush helpers also clobbered x1 (u-boot's void call
convention left mov x1, x0 in cache.S) which handed the kernel a wild
x0. The arguments ride in x20/x21 across the cache calls now, callee
saved, and the register setup is the last thing before the branch.
What is missing on purpose: no SMP bringup (secondary cores park),
no PSCI, no initrd or root filesystem handoff, single serial
console. Those come next.
Signed-off-by: Bradley Morgan <brads@mainlining.org>
|