Testing eBPF code: automatic tail call resolution

This post continues Testing eBPF code where it actually runs, which introduced the Arrange / Act / Assert test framework for eBPF programs written in Rust with Aya. If you haven’t read it, I recommend starting there: this post assumes you know what #[arrange], #[act] and #[assert] do. In the meantime, the framework moved out of Sarena into a repository of its own, and got a name: gladia.

This implementation: https://github.com/sarena-rs/gladia-ebpf

In the previous post, the act phase of a test ran the real production code. For example, if the production code contains:

#[classifier]
pub fn from_netdev(ctx: TcContext) -> i32 {
    ...
}

But how can we call this function from the test code? The answer is with tail calls. There are several steps needed to do this.

First, a ProgramArray is needed:

#[map(name = "entry_call_map")]
static entry_call_map: ProgramArray = ProgramArray::with_max_entries(10, 0);

This is defined in the eBPF test code. In this case, it creates a map with 10 entries.

The map then needs to be filled with the file descriptors (fd) of the programs that we want to call. For example, in slot 5 we can load the fd of the from_netdev program.

Typically, a userspace program, in this case the test runner, loads this array as follows:

let map = test_bpf
    .map_mut("entry_call_map")
    .ok_or_else(|| ...)?;
let mut program_array = ProgramArray::try_from(map)?;
let prog = bpf
    .program("from_netdev")
    .ok_or_else(|| ...)?
    .try_into()?;
let fd = prog.fd()?;    
program_array.set(FROM_NETDEV, fd, 0)?;

There are three parts that are important here:

  • entry_call_map - is the name of the map that resides in the test eBPF program (see above).
  • from_netdev - is the name of the program that we want to call. We take the file descriptor (fd) of this program.
  • FROM_NETDEV - is an integer constant that specifies the slot in which the fd of the from_netdev program is loaded.

This constant is defined in a shared crate, because the eBPF test program needs it for the tail call, and the userspace test runner needs it to fill the correct slot.

pub const FROM_NETDEV: u32 = 5;

Finally, the eBPF test code can do the tail call:

#[act(tc, "l2_announcement_arp_no_entry")]
pub fn l2_announcement_arp_no_entry_act(ctx: TcContext) -> TestStatus {
    unsafe {
        entry_call_map.tail_call(&ctx, FROM_NETDEV);
    }

    TestStatus::FrameworkError
}

Note that a successful tail_call does not return. For example, if the fd in slot 5 is invalid, the tail_call fails and TestStatus::FrameworkError is returned.

This is quite tedious, because there is information scattered across several crates.

In the eBPF test program (a separate crate):

entry_call_map.tail_call(&ctx, FROM_NETDEV);

In a shared crate (shared between the eBPF test program and the userspace test runner):

pub const FROM_NETDEV: u32 = 5;

And in the userspace test runner:

program_array.set(FROM_NETDEV, fd, 0)?;

Everything has to be kept in sync. If the wrong fd is loaded at the wrong index, either the wrong production program is tested, or the tail_call fails.

What if we could automate this?

The goal is that I only have to write this in an act program:

#[act(tc, "l2_announcement_arp_no_entry")]
pub fn l2_announcement_arp_no_entry_act(ctx: TcContext) -> TestStatus {
    tail_call!(&ctx, "from_netdev")
}

No constants, no wrapper, no list in the runner. The name of the production program is the only thing the writer of the test provides. Everything else — the program array, its size, the slot of each program, and loading the right programs in the test runner — follows from it. In other words, tail calls are resolved automatically.

The problem is that a tail call doesn’t take a name. It takes an index into a program array. So somewhere, someone has to collect all the names, give each one an index, size the map, and tell the runner which program goes where. So, the build has to do this.

Recall that, in essence, we needed three things:

  • a ProgramArray filled with file descriptors. Here, entry_call_map from the eBPF test program, and program_array from the userspace test runner are essentially the same thing.
  • a file descriptor (fd) of a program. Here we took the fd of the loaded eBPF program called from_netdev.
  • an index into the ProgramArray.

We could use several ProgramArrays, but in my setup I use only one. It could then look like this:

ProgramArray
index 0   --> fd of "something"
index 1   --> fd of "something_else"
index 3   --> fd of "another_thing"
index 4   --> fd of "to_netdev"
index 5   --> fd of "from_netdev"
...

The particular assignment of programs to indexes does not matter. What matters is that the eBPF test program and the userspace test runner agree on the same mapping.

So the build needs to do several things:

  • Find which production programs are used by the tests.
  • Assign each program an index and create the ProgramArray with the right size.
  • Generate the code.
  • Use the same mapping when loading the program fds in the userspace test runner.

Let’s start with the first one.

The first idea that came to mind is a procedural macro: make tail_call! a proc macro and let it do the work. That doesn’t work. A proc macro only sees the tokens it is applied to. The invocation in arp.rs has no idea that nat.rs also contains tail calls, so it can’t know how big the program array must be, or which index is still free.

What does see the whole crate is its build script. So the test crate gets a one-line build.rs:

fn main() {
    gladia::build_mapping();
}

build_mapping() walks the crate’s src/ directory, parses every .rs file with syn, and visits every macro invocation in it. Whenever it finds a tail_call!, it takes the second argument:

impl<'ast> Visit<'ast> for CallVisitor<'_> {
    fn visit_macro(&mut self, m: &'ast Macro) {
        let is_tail_call = m
            .path
            .segments
            .last()
            .is_some_and(|seg| seg.ident == TAIL_CALL_MACRO_NAME);

        if is_tail_call {
            // ...
            match &args[1] {
                Expr::Lit(ExprLit {
                    lit: Lit::Str(s), ..
                }) => self.calls.insert(s.value()),
                _ => fail("second argument must be a string literal"),
            };
        }

        syn::visit::visit_macro(self, m);
    }
}

The name has to be a string literal. That is a deliberate restriction: the set of tail calls must be known at build time, and a literal is the only thing a build script can read without running the code. Anything else — a variable, a constant, a missing argument — fails the build with the file and the offending call in the message.

The build script also tells cargo to rerun it whenever anything in src/ changes (cargo:rerun-if-changed on the directory). That gives a useful guarantee: whenever the test crate is built, the build script has seen every tail_call! in src/.

The calls end up in a BTreeSet, so they are unique and sorted. A name’s index is simply its position in that set:

0  from_netdev
1  to_host

Sorting matters more than it looks. read_dir returns files in no particular order, so without it the indices could differ from one machine — or one build — to the next. With it, the same sources always produce the same mapping, and a name used in ten tests still takes only one slot.

At build time, the build script scans the test crate for tail_call!, assigns each unique name an index, and generates the program array, the table and the macro

When the names and indices are known, the build script (i.e. build_mapping()) writes $OUT_DIR/__gladia_tail_calls.rs.

This generated file contains the following three constructs. Suppose the previous steps found two calls, from_netdev and to_host, and assigned indices 0 and 1 respectively. Then the following things are created.

First, the program array, sized to the number of unique calls. In this case 2:

#[aya_ebpf::macros::map(name = "__gladia_tail_call_map")]
pub static __gladia_tail_call_map: aya_ebpf::maps::ProgramArray =
    aya_ebpf::maps::ProgramArray::with_max_entries(2, 0);

Second, the tail_call! macro itself, with one arm per name:

#[macro_export]
macro_rules! tail_call {
    ($ctx:expr, "from_netdev") => {{
        let _ = unsafe { $crate::__gladia_tail_call_map.tail_call($ctx, 0u32) };
        gladia_ebpf::TestStatus::FrameworkError
    }};
    ($ctx:expr, "to_host") => {{
        let _ = unsafe { $crate::__gladia_tail_call_map.tail_call($ctx, 1u32) };
        gladia_ebpf::TestStatus::FrameworkError
    }};
    ($ctx:expr, $name:literal) => {{
        compile_error!(concat!("unknown tail call `", $name, "`; rebuild so the build script picks it up"));
        gladia_ebpf::TestStatus::FrameworkError
    }};
}

macro_rules! can match on a literal token, so tail_call!(&ctx, "from_netdev") compiles straight to unsafe { __gladia_tail_call_map.tail_call(ctx, 0) }, and tail_call!(&ctx, "to_host") compiles straight to unsafe { __gladia_tail_call_map.tail_call(ctx, 1) }. There is no lookup at runtime; the name exists only in the source.

Recall that a successful tail call never returns — execution continues in the production program — so the FrameworkError is only reached when the call fails, for example because the slot is empty. That is what lets an act program end with the macro, like in the example above.

The last arm is a catch-all. Under cargo it never matches: every name in the source has its own arm, because the build script ran first. It’s there for tools that look at stale generated code. rust-analyzer, for instance, doesn’t rerun build scripts when you edit a source file, and without the catch-all rust-analyzer would generate an error because none of the arms matches.

The test program now knows which index belongs to which name, but only as compiled code. The test runner, in userspace, needs the same information to fill the program array.

That’s the third generated item: a table in an ELF section of its own.

#[used]
#[unsafe(link_section = ".gladia_tail_call_section")]
pub static __gladia_TAIL_CALL_MAP: gladia_ebpf::TestEntryHeader<2> = gladia_ebpf::TestEntryHeader {
    version: 1u32,
    count: 2u32,
    size: core::mem::size_of::<gladia_ebpf::TestEntryCall>() as u32,
    entries: [
        gladia_ebpf::TestEntryCall { index: 0u32, name: *b"from_netdev\0\0\0..." },
        gladia_ebpf::TestEntryCall { index: 1u32, name: *b"to_host\0\0\0\0\0..." },
    ],
};

The layout is #[repr(C)] and shared between both sides: a header with a version, the size of one entry and the number of entries, followed by the entries — an index and a zero-padded, 64-byte name.

#[used] keeps the linker from throwing the table away, since no eBPF code ever reads it. The names are written as dereferenced byte strings (*b"..."), which is a [u8; 64] as long as the literal is exactly 64 bytes long; a name that doesn’t fit fails the build.

The generated file ($OUT_DIR/__gladia_tail_calls.rs) is pulled in with a macro at the top of the crate root.

#![no_std]

gladia_ebpf::include_generated!();

mod arp;
mod nat;

It has to come before the mod declarations: a macro_rules! macro is only visible in code that comes after its definition.

At test time, the test runner reads the .gladia_tail_call_section section from the test object with the object crate — no kernel involved — and checks the version and the entry size first. This means that a mismatch between the framework on both sides is reported instead of being misread. Then it wires everything up:

fn fill_entry_call_map(
    prod_bpf: &mut Ebpf,
    test_bpf: &mut Ebpf,
    entry_calls: &[(u32, String)],
) -> Res<()> {
    let map = test_bpf
        .map_mut(TAIL_CALL_MAP_NAME)
        .ok_or_else(|| TestRunnerError::MapNotFound(TAIL_CALL_MAP_NAME.to_owned()))?;
    let mut program_array = ProgramArray::try_from(map)?;

    for (slot, name) in entry_calls {
        load_bpf_program(prod_bpf, name)?;
        let fd = get_sched_classifier(prod_bpf, name)?.fd()?;
        program_array.set(*slot, fd, 0)?;
    }
    Ok(())
}

For every entry, it loads the production program with that name and puts its file descriptor (fd) in that slot (failing when the production program does not have that name). Only the programs the tests actually call are loaded; a production program that no test uses is never touched.

At test time, the runner reads the table from the test object, loads the listed production programs, and puts them in the program array that the act program tail calls into

The test runner doesn’t need a map (“program name” to index) from the user anymore. Now we can just write:

static PROGRAMS: &[u8] = aya::include_bytes_aligned!(concat!(env!("EBPF_OBJECTS"), "/ebpf-programs.o"));
static TEST_PROGRAMS: &[u8] = aya::include_bytes_aligned!(concat!(env!("EBPF_OBJECTS"), "/ebpf-test-programs.o"));

#[test]
#[ignore = "requires CAP_NET_ADMIN/CAP_SYS_ADMIN and a writable /run/netns"]
fn ebpf_test_runner() -> gladia::Res<()> {
    gladia::run_ebpf_test(PROGRAMS, TEST_PROGRAMS)
}

Note that the production code (in this case ebpf-programs.o) is as-is. It does not know it is being tested, and there is no dependency on gladia. The test program (in this case ebpf-test-programs.o) contains the resolved tail_calls (like tail_call(ctx, 0)) and the .gladia_tail_call_section table. The test runner (in run_ebpf_test) reads this section and fills the ProgramArray as described in the previous section.

The whole flow can now be reduced to two phases: build time and test time.

At build time:

  1. build.rs calls gladia::build_mapping(). The build script runs before the crate is compiled, scans the crate for tail_call! invocations, and collects their target program names. The names are sorted and deduplicated, and each unique name is assigned an index.
  2. build_mapping() generates $OUT_DIR/__gladia_tail_calls.rs. Among other things, it contains a tail_call! macro with an arm for each target name, resolving the name to its index, and a table containing the same name → index mapping for the userspace test runner.
  3. gladia_ebpf::include_generated!() is placed in lib.rs of the same test eBPF crate. It includes the generated tail_call! macro before the modules that use it, so tail_call!(&ctx, "from_netdev") resolves to the index assigned during the build.

At test time:

  1. The test runner reads the generated table from the test object’s .gladia_tail_call_section.
  2. For each entry, it loads the corresponding production program and obtains its file descriptor.
  3. It puts that file descriptor into the corresponding slot of the test program’s ProgramArray.
  4. When an act program executes tail_call!(&ctx, "from_netdev"), the kernel jumps to the real from_netdev production program.

So the name in the test source is the only piece of information that has to be written by hand. The build system derives the index and map size, and the test runner uses the generated table to put the right production program in the right slot.

The important part is that there is one source of truth: the tail_call! invocation in the test. There is no separate list of programs, manually maintained index constant, or name-to-index mapping for the test runner to keep in sync.

My first version of the build script generated one table per source file, each with the file’s name in its header. It seemed like useful information: which file calls what.

But in practice it bought nothing. The test runner needs index → name for every unique call, and that’s it.

The per-file tables stored a name once for every file that used it, needed a file name field that could overflow, and made the reader loop over a sequence of headers. The layout of those headers also depends on how the linker places them — which turned out to be one section per static, not one section with all of them merged.

One table for the whole crate is simpler on both sides, so that’s what remained.

  • Names must be string literals. That’s what makes build-time resolution possible, but it rules out choosing the target at runtime.
  • Names are at most 64 bytes, the size of the name field in the table.
  • Only direct invocations are found. syn does not parse the tokens inside other macros, so a tail_call! nested inside another macro’s arguments isn’t seen by the build script. Thanks to the catch-all arm, a tail_call! that isn’t found by the build script generates a compile error.
  • One include per object. The generated map and table must exist once, so include_generated!() belongs in one crate, typically the library of the test crate.
  • Stale editors. Until rust-analyzer reruns the build script (“Rebuild build scripts”), it doesn’t know about new names. The catch-all arm makes that obvious, but it doesn’t make it go away.
  • Still TC only, still one #[test]. The runner only handles TC programs, and all eBPF tests still run inside one Rust test; the summary makes the individual results visible, but cargo test can’t select them.

Tail calls were the one place where the framework still depended on bookkeeping by hand, and bookkeeping by hand is exactly what goes wrong silently.

Now the source code is the single source of truth: write tail_call!(&ctx, "from_netdev"), and the build script turns that name into a slot, a map of the right size, a macro arm and a table entry, and the runner turns the table entry into a loaded production program in the right slot.

Nothing to register, nothing to keep in sync — the tail calls resolve themselves, and the test still runs the real, verifier-approved production code in the kernel.

The next thing I want to implement is packets. Currently, an arrange program mostly copies prepared bytes into the packet — bytes generated with scapy, outside the test.

The packet builder in gladia-ebpf already tracks layers, and the plan is to let tests build their input layer by layer: Ethernet, IPv4 or IPv6, TCP or UDP, tunnels, with lengths and checksums filled in.

Its counterpart would be a packet verifier: build the expected packet in the assert program the same way, compare it with what the production code produced, and report the differences field by field — what the scapy integration does today, but in Rust, inside the framework, without Python.