Back to Blog

FreeBSD

Ding-Dong Ditch: The Doorbell Prank in FreeBSD's iflib

Ding-Dong Ditch: The Doorbell Prank in FreeBSD's iflib

Since 2017, FreeBSD®'s iflib(4) has been playing ding-dong ditch on network interfaces. When the transmit ring was nearly empty, iflib rang the interface's doorbell three times for every packet, and two of those rings announced nothing new. The interface still had to answer each one, and answering is slow. On a busy router, the cores spent most of their time on the doorstep instead of forwarding packets.

We found it on a Netgate® 4200, a small four-core router with 2.5 gigabit Ethernet ports. Flooded with the smallest packets, it topped out at 0.96 million packets per second (Mpps) in one direction, about a quarter of line rate. A one-line fix took it to 2.57 Mpps, and a follow-up efficiency change, which rings once per batch of packets, took it to 3.58 Mpps. That's within 4 % of the 3.72 Mpps the wire can carry.

This post follows the trail from the symptom to the prank, the fix, and what is still left.

How a router sends a packet

Every packet a router forwards ends with the same handoff. The operating system writes a short note about the packet (where it is in memory and how long it is) into a list the network interface reads from, called the transmit ring. It then writes a number into one of the interface's registers to say the list has grown. That write is the doorbell.

Ringing the doorbell is expensive. The register lives on the interface, not in ordinary memory, so each ring costs the processor hundreds of cycles, and the interface can only accept so many rings per second. Ringing less often saves a lot of work.

On FreeBSD, most Ethernet drivers share this machinery through iflib, a library that manages the transmit and receive rings. Each driver supplies only the hardware-specific parts, such as how to ring its doorbell. A bug in iflib therefore reaches every driver built on it, including igc, the driver for the 4200's Intel 2.5 GbE ports, which Netgate sponsored and adapted to iflib.

A closer look at iflib

Before iflib, each FreeBSD Ethernet driver carried its own ring management, buffer refill, interrupt handling and interface glue. Matt Macy wrote iflib to move that into one library; it was imported in May 2016 and shipped in FreeBSD 11.0. A driver supplies a handful of methods (encapsulate a packet into transmit descriptors, write the transmit doorbell, report completed descriptors, and count, fetch and refill receive descriptors), and iflib handles queue selection, the software transmit ring, interrupt and task scheduling, buffer recycling and netmap.

The software transmit ring, mp_ring, came from Navdeep Parhar's cxgbe driver. It is lock-free with many producers: whichever producer finds it idle becomes the consumer and drains it into the hardware ring. Each queue's receive and transmit work runs in a group task queue thread bound to one CPU core. With direct dispatch, the receive task carries a packet through the stack and transmits it before returning.

em, igb, igc, ix, ixl, ice, iavf, bnxt, vmx, axgbe, enetc, enic and mgb all use iflib. cxgbe, mlx5en and ena do not.

iflib milestones · from the FreeBSD source history

Date Change Commit
Dec 2014 mp_ring added to cxgbe (Navdeep Parhar, Chelsio) 7951040f8ada
May 2016 iflib imported (Matt Macy), in FreeBSD 11.0 4c7070db251a
Nov 2016 bnxt, the first iflib driver in the tree (Stephen Hurd) d933e97f9d7c
Jan 2017 em, igb and lem move to iflib (Sean Bruno), in FreeBSD 12.0 f2d6ace4a684
Mar 2017 Doorbell deferral by ring occupancy; the redundant writes begin 95246abb21da
Dec 2017 ix moves to iflib (Eric Joyner); ixl follows in Jun 2018, vmx in Jan 2019 c19c7afee3c8
Jul 2018 tx_abdicate: a separate task may drain the ring (Stephen Hurd) fe51d4cdfee0
Jan 2021 Doorbell check simplified; the commit records the design history 81be655266fa
Jul 2021 igc, for the I225 and I226 (Peter Grehan, sponsored by Netgate) 517904de5cca
Aug 2025 simple_tx, a transmit path without mp_ring (Andrew Gallatin) 84f8ca1bd11d
Oct 2026 Redundant doorbell writes removed 4a0ec469934a
Oct 2026 Batched doorbells proposed D60416

A router stuck at a quarter speed

To measure how fast the 4200 forwards, we sent it a steady stream of 64-byte frames (the smallest Ethernet allows), raised the rate step by step, and counted what came out the other side. Small packets are the hardest test, because the work per packet stays the same while the packets arrive faster.

With the router set to four queues per port, spreading the work over all four cores, forwarding stopped at 0.96 Mpps in one direction (unidirectional), and between 0.8 and 1.0 Mpps the cores went from 36 % busy to fully busy, far more than 25 % more traffic should cost. That ruled out the obvious explanation. If the router were simply doing too much work per packet then turning on the pf firewall, which adds work to every packet, would have lowered the ceiling. It did not. Letting the router take in more packets at a time did not raise it either. A profile of the cores pointed at a single function: the one that rings the doorbell.

Ringing with nobody at the door

iflib is meant to batch its doorbell rings. Before ringing, it counts the packets waiting to be announced and holds off until there are enough. How many is enough depends on how full the transmit ring is. When the ring is busy, iflib waits for a bigger batch. When it is nearly empty, iflib rings right away so a lone packet isn't left waiting.

The bug is in that last case. Below an eighth full, the required batch size is zero, and the rule was "ring if the number waiting is at least the number required". Zero is at least zero, so iflib rang even with nothing waiting, pointing the interface at a spot in the ring it had already reached.

iflib makes that check three times each time it sends: before it starts, after each packet, and at the end. A router that forwards each packet as soon as it arrives sends them one at a time into a nearly empty ring, so every packet paid for three rings. Two of them were for nothing. At 0.96 Mpps that is 2.9 million rings per second into one interface, about as many as the 4200's Intel I226 interface can accept. The cores were full because they were stuck on a doorbell that couldn't be answered any faster.

In the 4200's default setup, one queue per port, each port's traffic is handled by a single core, so in our one-way test one core did all the forwarding. It still paid three rings per packet, but it ran out of processing time at 0.78 Mpps, before the doorbell limit came into play. There the wasted rings don't cap throughput; they eat CPU time that could have forwarded more packets.

The code and how it got that way

A driver tells the NIC about new transmit descriptors by writing the ring's tail register, TDT on Intel parts. That is one uncached store to device memory, and on the I226 it costs the core about 400 cycles. The device accepts about 2.9 million of them per second.

iflib_txd_db_check() holds the doorbell until enough descriptors are pending, but the threshold scales with ring occupancy and is zero below an eighth full. The check before the fix, trimmed:

max = TXQ_MAX_DB_DEFERRED(txq, txq->ift_in_use);

/* force || threshold exceeded || at the edge of the ring */
if (ring || (txq->ift_db_pending >= max) ||
    (TXQ_AVAIL(txq) <= MAX_TX_DESC(ctx))) {
	dbval = txq->ift_npending ? txq->ift_npending : txq->ift_pidx;
	...
	ctx->isc_txd_flush(ctx->ifc_softc, txq->ift_id, dbval);
	txq->ift_db_pending = txq->ift_npending = 0;
	return (true);
}

With max at zero, ift_db_pending >= max is true when nothing is pending, and isc_txd_flush() writes a tail value the NIC already has. iflib_txq_drain() runs the check before its loop, after each packet and after the loop. Direct dispatch queues one packet and drains it at once, so the ring stays nearly empty and every packet paid three writes, two of them redundant.

Matt Macy's message for 81be655266fa records how the doorbell logic got there. The first version wrote the register every fourth packet and left stragglers to a callout. It could write a lower producer index after a higher one, which e1000 and ixgbe tolerated and which locked up ixl's MAC. A lock around the write fixed that and cost too much in contention. The third version, the code above, writes immediately on an empty queue and defers more as it fills. The same message notes that "an obvious missing optimization was to skip that doorbell write if db_pending is zero"; the note sat in the commit history for more than five years before anyone acted on it.

The extra rings date from March 2017 (95246abb21da) and first shipped in FreeBSD 12.0. The fix landed as commit 4a0ec469934a (FreeBSD review D60290), with a second commit, ac24108dabca (D60371), that tidies the function without changing its behaviour.

The fix

The fix (4a0ec469934a) is a single early return: if nothing is waiting, don't ring.

iflib writes each packet's note at the tail, rings the doorbell to say how far the ring now reaches, and the interface fetches from the head. Switch between the original code, where two of every three rings (in orange) arrive when the head already sits on the tail and announce nothing new, and the fix, which rings once per packet.

Now iflib skips both empty rings and rings once per packet, only when there is a new note to announce. The router forwards 2.57 Mpps in one direction, 2.7 times as fast, and at 0.8 Mpps its cores are 23 % busy instead of 36 %. Packets also wait less once the cores get busy: in a round-trip test on one queue, the original code's typical packet queued for 2 ms from 800 kpps per direction, while the fix still answered in about 130 µs.

One ring per batch

Packets leave one at a time, but they don't arrive that way. iflib takes up to 16 packets from the receive ring at once (its default receive budget, the same for every iflib driver) and hands them to the network stack in a single call. With direct dispatch, which the 4200 uses by default, each batch is forwarded before the next is collected.

A second change, now in FreeBSD review as D60416, holds the doorbell while a batch is being forwarded and rings once when the batch is done. It is controlled by a proposed knob, net.iflib.tx_db_batched, which is not yet part of FreeBSD. It only holds while the transmit ring is at most an eighth full, the range where iflib's own batching does nothing. Measured on one core, this cuts the count to 0.063 rings per packet, close to one per batch of 16. The charts call this version Batched. The review is also looking at how the change behaves when the machine is a TCP endpoint rather than a router.

Holding the doorbell delays the first packets of a batch. We measured it: a round trip through the router grows by 6 to 16 µs, and that is with a test sender that fills every batch; real traffic at low rates arrives singly and is rung at once. Under heavy load the effect reverses: with less time spent ringing, the slowest packets get through sooner.

All three versions. With batched doorbells, iflib writes the notes for up to 16 packets and rings once, and the interface fetches them all together.

Packets that leave outside a batch, such as those handed to a separate thread or sent by the router itself, still cost one ring each.

With both changes, the 4200 forwards 3.58 Mpps in one direction with its cores 70 % busy, and 5.41 Mpps in both directions at once (bidirectional), up from 1.54.

The results below compare the three versions in the same colours on every chart: Original (three rings per packet), Fix (one per packet) and Batched (one per batch). In the load chart, a line on the diagonal is forwarding everything it is offered and a line that goes flat has hit its limit; the lower panel shows how busy the cores were.

Four queues per port, pf disabled, one boot, versions switched by sysctl. At line rate one way, the router forwards 0.96 Mpps with the original code and 2.57 Mpps with the fix, both with the cores full, and 3.58 Mpps batched, at 70 % CPU. Both ways: 1.54, 4.24 and 5.41 Mpps. At 0.8 Mpps one way, all three forward everything, at 36, 23 and 20 % CPU.

How other drivers and systems batch

The proposed knob, net.iflib.tx_db_batched, exists only with D60416 applied, and is off by default there until the change has run on more hardware. The receive task hands the stack up to 16 frames per if_input() call, and with direct dispatch every forwarded frame is transmitted inside that call; each transmit queue used during the call holds its doorbell and rings once when the call returns.

mlx5en has held its send-queue doorbells during receive processing since 2022 (2d5e5a0d75b0). DragonFly's ifq staging and OpenBSD's transmit mitigation batch at the interface queue instead. Linux batches doorbells for XDP redirects; its regular forwarding path rings once per packet unless the qdisc has a backlog.

What each ring costs

To price a single ring, we had one core do all the forwarding and counted the cycles it spent per packet in three workloads: firewall off, firewall tracking 2,050 states, and firewall tracking about 200,000 states. Each ring removed saves 330 to 395 cycles, whatever the firewall is doing.

Removing two rings saves about the same number of cycles whether a packet takes 2,700 cycles to forward or 6,000: a big share of a cheap packet, a smaller share of an expensive one. With the firewall off, the fix speeds up a single core by 42 %; with 200,000 states tracked, by 14 %. Batching roughly doubles those gains.

A kernel counter counts the rings. Going from three rings to one saves 786, 668 and 734 cycles per packet in the three workloads; batching saves another 350–610.

+42 % with the firewall off, +18 % with 2,050 states tracked and +14 % with 198,657 states tracked. Batched: +74 %, +39 % and +29 %.

Where the time goes

A flame graph shows where the processors spent their time. Each bar is a function, its width is its share of all the samples, and the bars stacked on it are the functions it called. The three panels share one scale and were captured on all four cores at 2.0 Mpps offered one way, firewall off.

From one panel to the next, the doorbell bar shrinks and the idle bar grows.

Original: iflib_txd_db_check, drain_ring_lockless and bus_dmamap_load_mbuf_sg hold 74.5 % of all samples, and only 0.955 Mpps of the 2.0 offered gets through. Fix: iflib_txd_db_check falls to 13.8 % for its one ring per packet, everything is forwarded, and the cores are idle 41 % of the time. Batched: it falls to 0.9 % and idle rises to 52 %. Hover a frame to compare it across the panels; click to zoom; Escape resets.

What is left

In both directions, the batched version runs out of CPU at 5.41 Mpps, 73 % of line rate. Part of that limit is heat: above 5.6 Mpps offered, the cores reached 76 °C and throttled, so a cooler box would get somewhat further. No single function stands out the way the doorbell did. The largest is memcpy_erms, at 4.0 % of the time: iflib copies each received packet of up to 128 bytes into a fresh buffer so the original can stay in the receive ring. Next is m_free, at 3.6 %, freeing buffers once they have been sent.

In one direction, the router forwards 3.58 Mpps with its cores 70 % busy, so the CPU isn't the limit. The 3.7 % that doesn't get through at line rate is dropped while the interface's transmit ring is full, so the I226 itself is sending no faster than that. We haven't found what sets that limit.

Nobody runs off anymore

The prank is over. Since 4a0ec469934a, iflib rings only when something is waiting at the door. Batched doorbells improve that by waiting until a whole batch is on the step. Together they took the 4200's one-way forwarding from 0.96 to 3.58 Mpps. Because the fix is in iflib itself, every driver built on it stops ringing for nothing. No function in the profile now costs anything near what the doorbell did, so the next ceiling will have to be found somewhere other than the front porch.

Test setup

Router

Netgate 4200: four Intel Atom cores and Intel I226-V 2.5 GbE ports, driven by igc. With 64-byte frames, 2.5 GbE carries at most 3.72 Mpps in each direction; that is the line rate used throughout.

Traffic

A Netgate 4100 and a 6100 send and count the traffic with netmap pkt-gen, running at real-time priority: 64-byte frames carrying UDP, spread over 99,328 flows per direction. Ethernet flow control is off, so the router cannot slow the senders down.

Kernel

One pfSense development kernel for every figure: FreeBSD main as merged on 29 September 2026, plus the batching patch (D60416) and a counter of doorbell writes. The three versions are selected by sysctl within one boot:

  • Original: net.iflib.db_skip_empty=0, which restores the old behaviour.

  • Fix: the early return alone.

  • Batched: the fix plus the proposed net.iflib.tx_db_batched=1 from D60416.

The flame graphs were captured earlier, on the revision of the batching patch before its review changes and without the counter, with Original and Fix on one boot and Batched on another; every other figure uses the revised patch.

Load ramps

The forwarding and CPU chart. Four queues per port, direct dispatch (the receive task carries each packet through the stack and transmits it before returning), inline transmit, pf disabled. Each step offers an exact packet count over 10 seconds, and the receiving netmap application counts what arrives. CPU busy is the mean of the four cores over the step. In the batched ramp both ways, the cores were thermally throttled in the three steps from 5.6 Mpps offered up, at 76 °C.

Single core

The cycles and per-core charts. One queue per port, so one core forwards everything, with line rate offered in one direction. pf is off, or on with 1,024 flows (2,050 states) or 99,328 flows (198,657 states). Cycles per packet come from that core's cpu_clk_unhalted.core counter and doorbell writes from the kernel counter; each value is the mean of two rounds.

Flame graphs

Sampled with pmcstat -S cpu_clk_unhalted.core, recording full call chains on all four cores for 12 seconds at 2.0 Mpps offered one way, pf disabled. Folded and rendered with flamegraph.pl --hash, so a function keeps its colour across panels.