The Second Core (core1)
Stock QMK runs on one core and leaves the RP2040’s second core parked in the bootrom. PolyKybd starts it at boot and uses it for work that is pure computation: decompressing overlay images, rendering the Eden screensaver, and running DOOM.
The RP2040 has two cores. QMK runs on core0: USB, the matrix scan, the console, SPI to
the keycap displays and I2C to the status OLED. core1 runs a small service loop with no
RTOS thread, no console and no interrupts. core0 hands it jobs over the inter-core FIFO
and core1 computes into RAM. The CPU-heavy part of an overlay upload runs there, so a
burst of compressed keycap images does not stall the matrix scan.
flowchart LR C0["core0 — QMK main loop<br/>USB · matrix · SPI · I2C"] -->|"CORE1_CMD_* (FIFO)"| C1["core1 — service loop<br/>IRQs masked"] C1 -->|"writes"| OV["overlays[] (RAM)"] C1 -->|"writes"| EB["Eden keycap buffers (RAM)<br/>firmware 1.5.0+"] C1 -->|"done counter / sequence number"| C0
How core1 is started
Section titled “How core1 is started”- During boot,
core0launchescore1withmulticore_launch_core1()from thepolymod_core1community module (modules/polykybd/polymod_core1/). The splash moves from step 4 to step 5 when it returns. core1runs on its own static stack ofCORE1_STACK_SIZEbytes: 384 B before firmware 1.5.0, and 512 B from 1.5.0, the release that moves Eden ontocore1.- The first instruction
core1executes iscpsid i, andcore1_entry()(keyboards/polykybd/multicore_exec.c) masks interrupts again. They stay masked for the life ofcore1. ChibiOS’s handler for the core1 FIFO interrupt ends by raising an NMI, and that NMI runs the RTOS context switch on a core that has no thread state, which hungcore1.core1polls the FIFO, so it never needs an interrupt to wake. core1_entry()setsg_core1_entered. The boot breadcrumbs record it, so a crash record from a boot hang shows whethercore1had started.- The consequence: anything that needs an interrupt stays on
core0. That is SPI to the keycaps, USB, I2C and the console.core1only computes into RAM.
The commands
Section titled “The commands”| Command | Sent for | What core1 does | How core0 knows it is done |
|---|---|---|---|
CORE1_CMD_DECOMPRESS |
Compressed overlay upload (HID 16/17) |
RLE-decompresses one fragment into the overlay pool, and requests a display refresh when that overlay is on screen | core1_decomp_count catches up with core0_decomp_count |
CORE1_CMD_ROI_UPDATE |
Region overlay upload (HID 18/19) |
Copies a rectangle into an overlay | Same counter |
CORE1_CMD_RESET_BIT_IDX |
Start of a new upload | Resets the fragment position to 0 | Not tracked |
CORE1_CMD_EDEN_KEY + 1 argument word |
Eden idle screensaver (firmware 1.5.0+) | Computes one keycap’s plasma, ripple and comet pixels into a 360 B buffer | The job writes its sequence number |
CORE1_CMD_CRASH_TEST |
POLYKYBD_CRASH_TEST builds only |
Makes an unaligned store, to check that a core1 fault is recorded |
— |
While core1 is still working on a fragment, raw_hid_pre_receive_kb() leaves the next
HID packet in the USB driver’s queue (4 packets) for that main-loop pass, so the matrix
scan still runs. PRC overlays (HID 41) are decoded on core0, inside the HID
handler, behind the same gate.
Who owns core1, and when
Section titled “Who owns core1, and when”- The overlay service, from boot. This is the normal state.
- The Eden idle screensaver (firmware 1.5.0+). A keycap costs about 4.3 ms on
core0, and 93 % of that is computing the pixels.core1now computes them, two keycaps at a time, whilecore0collects each finished keycap, cuts the legend out and pushes it over SPI.- On the rig,
core0spends 40 ms per Eden frame instead of 115 ms, and Eden runs at 32 frames per 5 s instead of 25. - The status OLED’s idle screen draws every 75 ms instead of 150 ms while Eden runs
on
core1. - An overlay upload, a firmware write and a font-pack write all end the idle session
first, so they never share
core1with Eden. - If a keycap job does not finish within 100 ms, the rest of the session renders on
core0. Later sessions usecore1again only once it has finished every job it was handed, so a stalledcore1can never fill the FIFO and blockcore0.
- On the rig,
- DOOM.
doom_enter()resetscore1through the power-state machine and hands it to the engine, with a 4 KB stack carved from the overlay pool. The game renders oncore1alone;core0blits finished frames and passes keys through a ring buffer. Eden never usescore1while DOOM runs. On exit,core1is reset and the overlay service relaunched with a bounded handshake (100 ms, up to three tries). If that fails, overlay decompression stays off until the next reboot, but the keyboard keeps working.
Stack budget and diagnostics
Section titled “Stack budget and diagnostics”- The overlay and region jobs peak at about 164 B of
core1stack. The Eden keycap job is the deepest path: 300 B measured with-fstack-usage, plus about 96 B if a fault is taken at that depth. That is why the stack grew to 512 B with it. OPT_DEFS += -DCORE1_STACK_HWMinrules.mkfills the stack with a marker at launch and logs the deepest point reached. Pass it that way, not as-e EXTRAFLAGS=-DCORE1_STACK_HWM: a command-lineEXTRAFLAGSreplaces the flagsrules.mkadds, which drops the-Wcast-alignguard and breaks DOOM builds.- After adding a
core1command, measure the stack again and raiseCORE1_STACK_SIZEif the peak climbs.
Running your own job on core1
Section titled “Running your own job on core1”The pattern every existing job follows is in keyboards/polykybd/multicore_exec.c. A job
is a FIFO command word, optionally followed by one argument word, and a case in the
switch inside core1_entry().
What core1 may do
Section titled “What core1 may do”core1 has no RTOS thread, no console and interrupts masked. That rules out most of
QMK. A job may:
- read from flash and RAM, and write into RAM buffers that
core0is not touching at the same time; - call plain C functions that do not wait on an interrupt.
A job must not:
- call
uprintfor anything else that writes to the console; - touch SPI, I2C, USB, the shift registers or any other peripheral
core0drives; - call ChibiOS functions (
chThdSleep, mutexes, events). They need thread state thatcore1does not have; - write to flash.
core1’s code runs from XIP flash. While core0 erases or writes flash, it holds
core1 in reset (fw_staging_core1_lockout_begin()), so a job in flight at that moment
is lost. Font-pack and firmware writes, and overlay uploads, end the idle session first
for this reason.
Adding a command
Section titled “Adding a command”-
Add a command word to
fifo_command_t, continuing the0xcafe00NNseries. -
Publish the inputs before pushing the command. Write them into
volatilestatics, calldmb(), and then callmulticore_fifo_push_blocking(). The FIFO only carries 32-bit words, so anything bigger goes through shared memory. For one small argument, push it as a second word, ascore1_eden_key()does. -
Handle it in
core1_entry(). Read the argument withmulticore_fifo_pop_blocking()if you sent one. Do the work and write the result. Then calldmb(), and only after it publish completion by bumping a counter or writing a sequence number. Oncore0, read the completion value first, calldmb(), then read the result. A barrier only orders accesses on opposite sides of it, so admb()after the publish lets the compiler or the bus move a result store past the completion store. -
Never block
core0on the result.core0is the core that scans the key matrix, so a long wait there drops keystrokes. Use one of the two existing shapes:- Backpressure. Poll a “busy” predicate and defer new work until
core1catches up. The overlay path does this throughraw_hid_pre_receive_kb(), which leaves the next HID packet in the USB queue for one main-loop pass. - Deadline with a fallback. Give up after a fixed time and do the work on
core0. Eden falls back after 100 ms and does not usecore1again until it has drained every job it was given.
If
core0has to spin at all, wrap the spin incrash_phase_enter(CRASH_PHASE_CORE1_WAIT, …)/crash_phase_leave(). A hang then produces a crash record that says “core0 waiting on core1” instead of an anonymous watchdog reset. - Backpressure. Poll a “busy” predicate and defer new work until
-
Check that core1 is yours. DOOM takes
core1over completely. Gate your job on a predicate likecore1_eden_available(), which is true only oncecore1has reachedcore1_entry()and DOOM is not running. -
Add the
!USE_CORE1stub in the#elsebranch at the bottom of the file, so a build withoutcore1still links. Decide what the stub does: run the job oncore0, or report it as unavailable. -
Measure the stack. Build once with
OPT_DEFS += -DCORE1_STACK_HWMinrules.mkand raiseCORE1_STACK_SIZEif the new peak, plus about 96 B for a fault taken at that depth, no longer fits.
A minimal sketch of the shape (not code from the tree):
static volatile uint32_t my_job_seq; // written by core1 when donestatic volatile uint8_t my_job_out[64]; // result buffer
typedef enum { /* ...existing commands... */ CORE1_CMD_MY_JOB = 0xcafe0005, // followed by ONE argument word} fifo_command_t;
// in core1_entry()'s switch:case CORE1_CMD_MY_JOB: { uint32_t seq = multicore_fifo_pop_blocking(); compute_into(my_job_out, seq); // RAM only: no I/O, no printf dmb(); // result stores land first... my_job_seq = seq; // ...then completion is published} break;
// called from core0:void core1_my_job(uint32_t seq) { multicore_fifo_push_blocking(CORE1_CMD_MY_JOB); multicore_fifo_push_blocking(seq);}bool core1_my_job_done(uint32_t seq) { const bool done = (my_job_seq == seq); // read completion first... dmb(); // ...then the result reads may follow return done;}