Skip to content

genet MTU for upstream - #7617

Draft
nbuchwitz wants to merge 8 commits into
raspberrypi:rpi-6.18.yfrom
nbuchwitz:devel/genet-mtu-rpi
Draft

nbuchwitz wants to merge 8 commits into
raspberrypi:rpi-6.18.yfrom
nbuchwitz:devel/genet-mtu-rpi

Conversation

@nbuchwitz

Copy link
Copy Markdown
Contributor

Based on @6by9's #7614, with the goal of upstreaming the MTU support.

I've tested the original patch on CM4 and discovered some issues. So I've created a slightly different patch (series) which I intend to send to netdev. It also contains some fixes Sashiko would have flagged any way...

  1. TBUF_PKT_RDY_THLD (TBUF + 0x10) is never programmed and it stays at 0x80. At MTU 3824 TX iperf3 is stuck at 0.00 Mbit/s while ping works and the link is up. tx_pkts rises, but tx_good_pkts doesn't. Kudos to @wtschueller who discovered this Jumbo frame support on Pi4 ethernet (Genet) #5561
  2. UMAC_MAX_FRAME_LEN gets the MTU value, but it's a frame length and counts the FCS. Frames from 3824 up result in rx_length_errors (at least in my testing), so the real limit seems to be MTU 3806.
  3. Wire budget is THLD*16-2 = 3838, so 3824 + VLAN = 3842 breaks setups with VLANs configured. Therefore I used 3820.
  4. RX_BUF_LENGTH 10240 costs no throughput (936/941 at MTU 1500, same as unpatched) but is above KMALLOC_MAX_CACHE_SIZE on arm64, thus it cant hurt to derive it from the MTU instead.

0xf0 seems to be the real limit: 0xfb receives fine but resulted in TX hard-hung on my setup.

Happy to add @6by9 as Co-developed-by since it's based on your findings. But this requires a Signed-off, which I wouldn't add without consent.

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

39dfaf2 contains a brutal approach to make MTU 9000 work (without any offloading). Performs quite ok, but needs more testing

@6by9

6by9 commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

I'm still waiting on documentation from Broadcom to read the official word on how jumbo frames with offload was meant to work (if it was).

Seeing as it was the offloading headers that seemed to cause issues, I did wonder if disabling offloading would allow it to work with bigger buffers. I only had a very quick read through the patches, but wonder if we can "dynamically" disable offload when the mtu is increased above the magic threshold. Possibly not based on the comment of losing the queues as well.

I had considered VLAN headers, but didn't know the answer off the top of my head, and wasn't in a position to set up VLANs to test. Thanks for taking care of it.

I'm not fussed over Co-developed-by:. I'm very grateful that someone else is having a look at the patches, particularly when they're looking to upstream it too.

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

@ffainelli and @Ryceancurry if you can spare some time, your thoughts on this would be really appreciated (as always). Thanks!

@starchivore

Copy link
Copy Markdown

https://lore.kernel.org/netdev/20260406-devel-autonomous-eee-v1-1-b335e7143711@tipi-net.de/t/

Other BCM54xx PHYs likely have the same AutogrEEEn register layout, but I only have access to the BCM54210PE/BCM54213PE datasheets.


https://datasheets.raspberrypi.com/cm4/cm4-datasheet.pdf#page=7

The CM4 has an on-board Gigabit Ethernet PHY — the Broadcom BCM54210PE

https://www.broadcom.com/products/ethernet-connectivity/phy-and-poe/copper/gigabit/bcm54210

• Supports jumbo packets up to 18 KB


https://magazine.raspberrypi.com/articles/raspberry-pi-4-in-detail

The BCM54213PE chip connects the Ethernet to a high-speed interface to the CPU.

https://www.broadcom.com/products/ethernet-connectivity/phy-and-poe/copper/gigabit/bcm54213

• Support for jumbo packets up to 10 KB


While we do understand the importance of taking one step at a time, it would be great to test whether 10K (BCM54213PE) and 18K (BCM54210PE) are genuinely supported by the hardware or otherwise. Thanks.

@Ryceancurry

Ryceancurry commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

With the status blocks off jumbo frames seem to come through.

Can you give me more color on the failure? Do we see fragmented packets? Is the packet corrupted? Or do we not receive a RX descriptor at all?

Full disclosure, I threw AI at the RTL(I'm a SW guy), it suggests a RTL bug where the RSB is reserved at every packet ready threshold. So I wonder if we are seeing a 64B hole between each 3820B chunk within the jumbo packet. At least that is the running theory right now. I will continue to dig.

@herisson-88

Copy link
Copy Markdown

Status blocks off, what's the true max frame length the MAC can handle ?

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

Thanks for looking into this too!

With the status blocks off jumbo frames seem to come through.

Can you give me more color on the failure? Do we see fragmented packets? Is the packet corrupted? Or do we not receive a RX descriptor at all?

With status block enabled and threshold at 0xf0 I get a descriptor (one per oversized frame):

desc len=3904 status=0x0f402000 SOP=1 EOP=0

3904 = 64 (RSB) + 2 (align) + 3838. Payload is fine and matches my test pattern. I also don't see any holes, just a hart cut off.

I've also tested with 3840, 5000 and 9014 B frames and all of them produce the same descriptor.

Some things I've noticed and might be worth mentioning:

  • the MAC MIB counts the frame correctly (9014 B increments rx_4096_9216_oct), so the MAC gets the complete frame and I suspect the loss somewhere in RBUF to RDMA handoff
  • no corruption to follow-up traffic, sending 9014 and 1514 B one after the other, every 1514 one is OK

For contrast, with RBUF_64B_EN and TBUF_64B_EN cleared, MTU 9000 works at line speed with byte exact payloads and the threshold still at 0xf0.

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

Seeing as it was the offloading headers that seemed to cause issues, I did wonder if disabling offloading would allow it to work with bigger buffers. I only had a very quick read through the patches, but wonder if we can "dynamically" disable offload when the mtu is increased above the magic threshold. Possibly not based on the comment of losing the queues as well.

I've tested further and came up with a solution which allows to switch to higher MTU on a live interface (tested 1514, 4096, 8192, 9014 B frames with threshold at 0xf0). Anyway, blocks needs to be disabled for anything higher. If I keep the TSB to preserve TX checksum offload, RX is still fine at 986 Mbit/s but TX drops to 0.

A while ago I proposed to get rid of the TX queues in genet [1]. Florian and Justin reviewed and tested it, but the reasoning was not good enough. Even though the queues are not absolutely blocking it,the TSB has no queue selection role anymore and it would simplify the jumbo patch. Might be worth a v2.

[1] https://lore.kernel.org/netdev/20260612205915.3156127-1-nb@tipi-net.de/

@herisson-88

Copy link
Copy Markdown

Patch tested on Audiolinux.
Many hours of Diretta streaming at MTU 9000 without issue.

@Ryceancurry

Ryceancurry commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

Status blocks off, what's the true max frame length the MAC can handle ?

As far as I can see the only limitation is the MAC's 14 bit frame len field. so 16383B

I reproduced the 9000B frames with RSB enabled. I printed out the entire 9000B packet and see corruption at each PKT RDY THRESHOLD. 3838B and ~7700B. This confirms my suspicion. The HW puts a 64B header per PKT RDY THRESHOLD. I think the correct way to do this is to set rx_buf_size to PKT RDY THRESHOLD. Then use rx scatter gather with multiple descriptors. We need to strip 64B off of each fragment. Unfortunately this means a big rework on the RX side.

@herisson-88

Copy link
Copy Markdown

@nbuchwitz we can test 16K with the current patch, it is just a question of ENET_MAX_JUMBO_MTU ?

@nbuchwitz
nbuchwitz force-pushed the devel/genet-mtu-rpi branch 2 times, most recently from b060753 to 124162d Compare September 12, 2026 18:57
@nbuchwitz

Copy link
Copy Markdown
Contributor Author

I reproduced the 9000B frames with RSB enabled. I printed out the entire 9000B packet and see corruption at each PKT RDY THRESHOLD. 3838B and ~7700B. This confirms my suspicion. The HW puts a 64B header per PKT RDY THRESHOLD. I think the correct way to do this is to set rx_buf_size to PKT RDY THRESHOLD. Then use rx scatter gather with multiple descriptors. We need to strip 64B off of each fragment. Unfortunately this means a big rework on the RX side.

That helped a lot, thanks. I swapped the MTU 9000 patch for your approach and it works well. The TSB even can stay on with a little quirk. I also bumped max_mtu to what the 14 bit UMAC_MAX_FRAME_LEN allows, 16347.

Do you now if the status block bug is in all GENET (non v1) versions? I only have v5 here to test.

@herisson-88

Copy link
Copy Markdown

@nbuchwitz work great in 1G (I get 9184 the limitation is on other side) but configured in 100M there are packet loss with mtu > 9080

@antonellocaroli

Copy link
Copy Markdown

I reproduced the 9000B frames with RSB enabled. I printed out the entire 9000B packet and see corruption at each PKT RDY THRESHOLD. 3838B and ~7700B. This confirms my suspicion. The HW puts a 64B header per PKT RDY THRESHOLD. I think the correct way to do this is to set rx_buf_size to PKT RDY THRESHOLD. Then use rx scatter gather with multiple descriptors. We need to strip 64B off of each fragment. Unfortunately this means a big rework on the RX side.

That helped a lot, thanks. I swapped the MTU 9000 patch for your approach and it works well. The TSB even can stay on with a little quirk. I also bumped max_mtu to what the 14 bit UMAC_MAX_FRAME_LEN allows, 16347.

Do you now if the status block bug is in all GENET (non v1) versions? I only have v5 here to test.

Could you show me the little quirk you used to keep the TSB enabled? On my Raspberry Pi 4 / GENET v5, MTU 13500 works, but around 13505 it becomes unstable and MTU 14000 fails with RX CRC errors. I noticed that a 14000 MTU results in a 14014-byte skb becoming a 14078-byte DMA buffer after the 64-byte TSB is added.

@nbuchwitz

nbuchwitz commented Sep 13, 2026

Copy link
Copy Markdown
Contributor Author

Could you show me the little quirk you used to keep the TSB enabled? On my Raspberry Pi 4 / GENET v5, MTU 13500 works, but around 13505 it becomes unstable and MTU 14000 fails with RX CRC errors. I noticed that a 14000 MTU results in a 14014-byte skb becoming a 14078-byte DMA buffer after the 64-byte TSB is added.

The "quirk" is to dynamically switch of TX checksum based on the mtu (threshold is the previous 3820). See the last patch for details.

Pi4 is afaik limited by the phy around 10k (see comment above). Cm4 should (theoretically) something around 18k

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

@nbuchwitz work great in 1G (I get 9184 the limitation is on other side) but configured in 100M there are packet loss with mtu > 9080

Haven't tested it yet with fast ethernet. If the time permits I will do some measurements with different mtu and speed. I want to measure the cpu impact of sw checksum. For jumbo frames I assume not much of a penalty

@antonellocaroli

Copy link
Copy Markdown

@nbuchwitz work great in 1G (I get 9184 the limitation is on other side) but configured in 100M there are packet loss with mtu > 9080

Haven't tested it yet with fast ethernet. If the time permits I will do some measurements with different mtu and speed. I want to measure the cpu impact of sw checksum. For jumbo frames I assume not much of a penalty

Thanks, that clarifies the TSB quirk. Interestingly, with two Pi4 Model B (Rev 1.1 and Rev 1.5) directly connected, I can get MTU 13500 working reliably in one direction (10/10 pings), while the opposite direction fails. Around 13503–13507 it becomes unstable/fails. So the Pi4 PHY seems capable of going significantly beyond 10k in at least some cases. Do you know what exactly imposes the ~10k PHY limit you mentioned (PHY register/buffer/specification), and whether it differs between Pi4 board revisions?

@nbuchwitz

nbuchwitz commented Sep 13, 2026

Copy link
Copy Markdown
Contributor Author

Do you really see such big payload or is this already capped by phy and it just works "magically" with the 10k limit?

Limit is stated in the datasheet. So I'd assume it's related to the buffer / state machine

@antonellocaroli

Copy link
Copy Markdown

Do you really see such big payload or is this already capped by phy and it just works "magically" with the 10k limit?

Limit is stated in the datasheet. So I'd assume it's related to the buffer / state machine

Yes, at least at the GENET MAC/driver level I really see the full size. For 10 successful MTU 13500 pings, txq3_packets increases by 10 and txq3_bytes by 135140, i.e. exactly 13514 bytes per packet. tx_oversize also increases by 10 on TX and rx_oversize by 10 on RX, with no additional CRC errors. I'm using ping -M do, so there is no IP fragmentation.

However, I haven't verified on the wire between MAC and PHY, so you're right that this doesn't prove the PHY actually handles the full ~13.5K frame as such. Interestingly, Rev 1.1 -> Rev 1.5 works at MTU 13500, while Rev 1.5 -> Rev 1.1 fails, even though the receiving side counts the request and generates a 13514-byte reply.

Which PHY datasheet/section states the ~10K limit? I'd like to check exactly what that limit refers to.

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

@antonellocaroli

antonellocaroli commented Sep 13, 2026

Copy link
Copy Markdown

https://www.broadcom.com/products/ethernet-connectivity/phy-and-poe/copper/gigabit/bcm54213pe

Support for jumbo packets up to 10 KB

Thanks!

yes, I confirmed that both of my Pi4s (Rev 1.1 and Rev 1.5) are using the BCM54213PE PHY (phy_id 0x600d84a2).

So the 10 KB limit you mentioned is indeed the one stated in the BCM54213PE datasheet.

However, MTU 13500 is really passing end-to-end in my tests: with ping -M do -s 13472, I can get 10/10 replies with no fragmentation. So it looks like the 10 KB figure is a guaranteed/specification limit rather than a strict hardware cutoff.

Above that it becomes unreliable very quickly (around 13503–13507 in my tests), and at MTU 14000 it fails. I also see RX CRC errors when operating around this boundary, so this is clearly outside the PHY's guaranteed operating range.

Interestingly, both Pi4 revisions use exactly the same BCM54213PE, so the different behaviour I saw between the two boards isn't explained by a different PHY model.

I agree that this could be related to an internal PHY buffer/state-machine limit rather than a simple hard packet-size check.

@herisson-88

Copy link
Copy Markdown

I will be able to test 16k CM4 tomorrow

bcmgenet_hfb_init() runs INIT_LIST_HEAD() on priv->rxnfc_list, which drops
every rule off the list, and bcmgenet_open() calls it on each ifup. Every
rule the user configured is silently lost:

  # ethtool -N eth0 flow-type ether dst $MAC action 0
  Added rule with ID 0
  # ethtool -n eth0 | grep -c Filter:
  1
  # ip link set eth0 down && ip link set eth0 up
  # ethtool -n eth0 | grep -c Filter:
  0

Initialise the lists once at probe and restore the rules on open, as
bcmgenet_resume() already does.

Fixes: 3e37095 ("net: bcmgenet: add support for ethtool rxnfc flows")
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
bcmgenet_netif_stop() disables the Tx queues first and stops Tx NAPI
several steps later. A completion in flight calls netif_tx_wake_queue() in
between, so a queue runs again while bcmgenet_dma_teardown() and
bcmgenet_fini_dma() free the rings, and a transmit entering that window
touches freed control blocks.

Close is safe because dev_deactivate_many() stops the qdisc before
ndo_stop() runs. bcmgenet_suspend() leaves it running, so stop Tx NAPI
before the queues.

Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
bcmgenet_netif_stop() already takes stop_phy. Give the start side the same
choice so a caller that left the PHY running can bring the datapath back
without tripping the phy_start() state check.

No functional change.

Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
ENET_MAX_MTU_SIZE holds a frame length, not an MTU. Both users program it
into hardware that wants a frame length, so the name misleads as soon as
the MTU stops being fixed at ETH_DATA_LEN. Name the receive offset too,
which is open coded as 66.

No functional change.

Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
The receive buffer is a fixed 2048 bytes and the packet ready thresholds
keep whatever the reset left them at, so neither follows the MTU.

Compute the threshold from the MTU, program it into RBUF and TBUF, and
size the buffer to what that threshold lets the hardware deliver, the
status block on top of the threshold itself. The MTU is still fixed at
ETH_DATA_LEN, so the threshold works out as the reset default and only the
buffer grows, by the 64 bytes of status block it always had to hold.

Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
The driver never sets dev->max_mtu, so the MTU is stuck at ETH_DATA_LEN.

The thresholds follow the MTU, and their registers are 8 bit in units of
16 bytes and want a multiple of the 256 byte burst size, so 0xf0 is the
largest usable value. That leaves an MTU of 3820 once the alignment bytes,
the Ethernet header and a VLAN tag are taken off.

Resize the buffers and rewrite the registers in place, so the PHY keeps
running and the link stays up. A failed allocation falls back to the
previous size, and if even that fails take the interface down rather than
run on rings that are not there.

Link: raspberrypi#5561
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
@nbuchwitz
nbuchwitz force-pushed the devel/genet-mtu-rpi branch 2 times, most recently from 7ad4826 to 2121751 Compare September 13, 2026 19:28
A frame longer than the packet ready threshold is not truncated. The
hardware splits it across descriptors and writes a status block at the
start of each one, so the first fragment arrives with SOP set and no EOP
and is dropped as fragmented. That caps the MTU at 3820.

In order to support a larger MTU, the fragments have to be reassembled
after the status blocks have been stripped.

The MAC only checksums frames up to the threshold and drops longer ones
silently, so check those in software. That costs little at jumbo sizes,
where the larger frame saves more per packet overhead.

Use 16347 as the maximum MTU, based on the 14 bit UMAC_MAX_FRAME_LEN,
which counts the FCS.

Suggested-by: Justin Chen <justin.chen@broadcom.com>
Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
@nbuchwitz

Copy link
Copy Markdown
Contributor Author

I've created a second draft PR so it can be A/B tested. The other PR contains the patch set + 3 prep patches and is based on the series for net-next, as the driver upstream has been converted to page pool.

#7623

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

@nbuchwitz work great in 1G (I get 9184 the limitation is on other side) but configured in 100M there are packet loss with mtu > 9080

Haven't tested it yet with fast ethernet. If the time permits I will do some measurements with different mtu and speed. I want to measure the cpu impact of sw checksum. For jumbo frames I assume not much of a penalty

I can reproduce the issue and also found the likely culprint. The phy driver needs to call bcm_phy_enable_jumbo() like it's already done in bcm7xxx.c. I will add a patch for this.

@herisson-88

herisson-88 commented Sep 14, 2026

Copy link
Copy Markdown

Test report (relayed from Snyder, AudioLinux)

Kernel: head commit 124162d (2026-09-12), base rpi-6.18.y — i.e. the reassembly version, before the bcm_phy_enable_jumbo() fix. Hardware: Pi 4. MTU set to 16000 on the target.

Method: 5000 consecutive ping -c 1 -w 1 -M do -s <size>, stop at first failure, size increased step by step. Repeated at three link speeds.

Results:

  • 10 Mb/s: stable up to MTU 16000, no loss.
  • 100 Mb/s: loss starts at 9024 bytes.
  • 1 Gb/s: loss starts at 13478 bytes.

The threshold is not sharp. Loss begins at roughly 1–2 packets per 1000, then grows gradually with packet size until nothing passes. The exact point drifts over time, with no traffic other than the pings.

For reference, the same 100 Mb/s wall was seen earlier here at 9080 bytes, and 13500 at 1 Gb/s by antonellocaroli on Pi 4 — three independent setups, same two ceilings within a few tens of bytes.

The gradual, drifting, speed-dependent behaviour looks consistent with the PHY elastic FIFO not being enabled rather than a MAC limit. We'll rerun the same script once the bcm_phy_enable_jumbo() patch is in the branch and report whether the thresholds move.

MTU 9000 on the same kernel is reliable at all three speeds and covers everything the Diretta use case needs (up to DSD512 / 32-bit 705.6 kHz).

Will test this on CM4 with real streaming this evening.

Jumbo packets need two bits that default to off, extended packet length in
the auxiliary control register and PCS transmit FIFO elasticity in the
extended control register. The latter raises the transmit limit from 4.5
KB to 9 KB at the cost of 16 ns of 1000BASE-T transmit latency, and the
two together take copper mode to 10 KB.

Without them such frames are lost on a 100M link while the same frames
pass at 1G. On a Raspberry Pi CM5, which uses a BCM54210PE, 9142 byte
frames at 100M are lost 20 out of 20 with the MAC counting every one
as transmitted. The same frames over the same path at 1G arrive intact.

Set both, which bcm_phy_enable_jumbo() already does for bcm7xxx. The
frames then arrive and the payloads check out byte for byte.

Signed-off-by: Nicolai Buchwitz <nb@tipi-net.de>
@nbuchwitz

nbuchwitz commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author

79b69ad should fix this hopefully. Please note if you're testing against another RPi you might want to patch both systems, as this affects the transceiver side.

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

@herisson-88 It would be also great if you can test #7623 as this will be the version for upstream (with backports). It also contains the phy fix

@herisson-88

Copy link
Copy Markdown

@nbuchwitz currently compile kernel with #7623, RTL8156BG arrived (hope there is not issue on this side to not pollute with false positive... It is supposed go to 16K - friends are using it for that.) I have also another old RPI4 that I could use with the CM4 in the other side to use your kernel patchs on both side.

@Ryceancurry

Copy link
Copy Markdown
Contributor

Do you now if the status block bug is in all GENET (non v1) versions? I only have v5 here to test.

Maybe the right way to think about it is that the original one descriptor per jumbo frame wasn't meant to work. It just so happened to work with the RSB disabled. So it is safe to assume this is the correct way to do things in all revisions of genet. The RTL also corroborates.

Thanks for taking this on. Good work so far!

@herisson-88

herisson-88 commented Sep 14, 2026

Copy link
Copy Markdown

bcmgenet jumbo patch on CM4 — native DSD over jumbo frames, test report (2026-09-14)

Three tests, all conclusive: a native DSD512 Diretta stream (2 × 22.58 MHz, ~46 Mb/s) carried in
~16 KB Ethernet frames into the CM4 at 1 Gb/s, sustained, zero errors; a ping test at
10 / 100 / 1000 Mb/s with 9000 to 16000-byte frames, zero loss; and native DSD256 at 100 Mb/s
with 10000-byte frames — the exact case that failed on 6.18.50-1 — streaming clean.

Setup

  • Target: Holo Audio Red = Raspberry Pi Compute Module 4 Rev 1.1, AudioLinux (Arch Linux ARM),
    linux-rpi 6.18.50-2 PREEMPT_RT (Piero's package of 2026-09-14, build
    Mon Sep 14 13:35:35 CEST 2026) with your bcmgenet jumbo patch, GENET 5.0, external RGMII,
    interface end0. Diretta Target library 148 (diretta_alsa_target), ExtEtherMTU set to
    the link MTU of each test.

  • Host: x86 NUC, Fedora 44 (kernel 7.2.4), USB 3 2.5 GbE dongle Realtek RTL8156B (in-tree
    r8152, maxmtu 16362), hardware offloads off. Diretta host: DirettaRendererUPnP, transfer
    mode VarMax with a 20 ms cycle, DSD sent natively (no DoP).

  • Direct cable, no switch. The CM4 is 1 GbE, so 1000 Mb/s is the maximum link speed.

    CM4

    ip -d link show end0 → mtu 16000 minmtu 68 maxmtu 16347

    host

    ip -d link show eth-diretta2 → mtu 16000 minmtu 68 maxmtu 16362

1. Native DSD512 at 1000 Mb/s, MTU 16000 on both ends — 30 s of playback

Host side, interface counters (/sys/class/net/eth-diretta2/statistics):

Frames sent 10 948 in 30 s (364 per second)
Average frame size on the wire 15 979 bytes
Throughput 46.65 Mb/s
Host tx_errors / rx_errors 0 / 0

CM4 side, ethtool -S end0 before and after the same 30 s:

Counter Before After Delta
rx_pkts 550 751 561 709 +10 958 (matches the host)
rx_crc_errors 0 0 0
rx_frame_errors 0 0 0
rx_jabber 0 0 0
rx_length_errors 0 0 0
rx_oversize 93 921 104 858 +10 937 (see the note at the end)
rx_good_pkts 456 830 456 851 +21 (the small control frames only)

Load average on the CM4 during playback: 0.32. Audio continuous, no dropout during the session;
the listener's verdict was "perfect".

For scale: with the Diretta host SDK's default Auto transfer mode, a 48 kHz PCM stream on the
same link was measured at ~784-byte frames, 500 per second. Large frames only appear once the
transfer mode asks for them (VarMax); with it, DSD512 travels in 364 frames per second.

2. Ping test: 12 combinations, zero loss

Method: 100 ICMPv6 echo requests per combination, DF bit set (no fragmentation allowed),
20 ms apart, run from the host to the CM4 and then from the CM4 to the host. Link speed forced
on the host side, the CM4 following by auto-negotiation. Sizes are IP packet sizes.

Link speed 9000 bytes 10000 bytes 14000 bytes 16000 bytes
1000 Mb/s 0 lost / 100 0 lost / 100 0 lost / 100 0 lost / 100
100 Mb/s 0 lost / 100 0 lost / 100 0 lost / 100 0 lost / 100
10 Mb/s 0 lost / 100 0 lost / 100 0 lost / 100 0 lost / 100

Identical result in the other direction (CM4 → host): 0 lost out of 100 in every combination.
Total 2400 packets sent, 2400 received. CM4 counters rx_crc_errors, rx_frame_errors and
rx_jabber stayed at 0 throughout.

This closes the 100BASE-TX issue of the 2026-09-13 report (frames above ~9080 bytes were
dropped with jabber/CRC errors on 6.18.50-1): with the PHY fix in 6.18.50-2, 16000-byte
frames pass at 100 Mb/s and even at 10 Mb/s.

3. Native DSD256 at 100 Mb/s, MTU 10000 on both ends — the case that failed on 6.18.50-1

Link forced to 100 Mb/s full duplex, ExtEtherMTU=10000, native DSD256 (2 × 11.29 MHz,
~22.6 Mb/s), 30 s of playback:

Frames host → CM4 8 481 in 30 s (282 per second), average 10 004 bytes
Throughput / link load 22.6 Mb/s, 22 % of the 100 Mb/s link
CM4 rx_pkts +8 763 (audio frames + control)
CM4 rx_crc_errors / rx_frame_errors / rx_jabber 0 / 0 / 0
Renderer warnings none; audio plays perfectly

Note on the MIB statistics (cosmetic, no effect on traffic)

Measured on both the 16000-byte and the 10000-byte streams: every jumbo frame increments
rx_pkts and rx_oversize by one, rx_good_pkts does not move, and no size bucket moves either
(the last one is rx_4096_9216_oct). Example over 10 s of the 10000-byte stream: rx_pkts
+2 826, rx_oversize +2 826, rx_good_pkts +0, host sent 2 827 frames. The frames are delivered
intact (CRC 0, the DAC plays them); only ethtool -S reads as if every jumbo frame were bad.
Whether that is a threshold register worth updating along with the MTU, or simply how the GENET
MIB block is defined, is your call.

Reproduction

# both ends
ip link set <if> mtu 16000
# CM4 counters, before and after a stream
ethtool -S end0 | grep -E 'rx_pkts|rx_good_pkts|rx_oversize|rx_crc_errors|rx_frame_errors|rx_jabber'
# host frame rate / size
cat /sys/class/net/<if>/statistics/tx_packets /sys/class/net/<if>/statistics/tx_bytes
# ping test, DF bit, IP size S
ping -6 -M do -s $((S-48)) -c 100 -i 0.02 <link-local>%<if>

Thanks for the patch — it does exactly what was hoped for.

Will test with Rpi 4 tomorow, but CM4 is 16k compliant now.

@herisson-88

Copy link
Copy Markdown

bcmgenet jumbo patch on Raspberry Pi 4 Model B (BCM54213PE) — test report (2026-09-15)

Follow-up to the CM4 report above, same kernel build, this time with the Pi 4 Model B as the
Diretta host. Two results: a ping sweep at 1 Gb/s from 9020 to 16000-byte frames, 10 000 pings,
zero loss — the ~13 478-byte wall seen on Pi 4 before the bcm_phy_enable_jumbo() fix is gone;
and a native DSD512 album playing right now through ~16 KB frames from the Pi 4 to the CM4,
16 minutes and 8 tracks so far, frames counted on both ends, zero errors.

Setup

  • Host: Raspberry Pi 4 Model B Rev 1.4 (revision code b03114, 2 GB, Sony UK), BCM2711
    (4× Cortex-A72 r0p3, 1.8 GHz, arm_boost=1), GENET v5, external RGMII, PHY BCM54213PE
    (phy_id 0x600d84a2, the one whose datasheet says "jumbo packets up to 10 KB"), interface
    end0. AudioLinux (Arch Linux ARM), linux-rpi 6.18.50-2 PREEMPT_RT (Piero's package of
    2026-09-14, build Mon Sep 14 13:35:35 CEST 2026, i.e. the reassembly version with the
    PHY jumbo fix). Bootloader EEPROM 2020-09-03, start4.elf 2026-09-14.

  • Target: Holo Audio Red = Raspberry Pi Compute Module 4 Rev 1.1, same kernel package
    (see previous report), ExtEtherMTU=16000.

  • Direct cable between the two GENETs, no switch, autoneg on, link at 1000 Mb/s full duplex.
    Both sides patched, as advised for the PHY-side fix.

  • Diretta host software: DirettaRendererUPnP 2.5.19, --mtu 16000, transfer mode VarMax,
    DSD sent natively (no DoP). The Pi 4's LAN side is a separate USB 3 GbE dongle (ASIX AX88179A,
    ax88179_178a, MTU 1500) so end0 carries Diretta traffic only.

    Pi 4

    ip -d link show end0 → mtu 16000 minmtu 68 maxmtu 16347

    CM4

    ip -d link show end0 → mtu 16000 minmtu 68 maxmtu 16347

1. Ping sweep at 1000 Mb/s, MTU 16000 on both ends — 1000 pings per size

ping -6 -M do -c 1000 -i 0.005 -W 1 -s <payload> <link-local>%end0, Pi 4 → CM4:

Frame (IP size) Payload Sent / received Loss RTT min / avg / max (ms)
9 020 8 972 1000 / 1000 0 % 0.224 / 0.241 / 0.544
10 020 9 972 1000 / 1000 0 % 0.244 / 0.263 / 0.696
11 020 10 972 1000 / 1000 0 % 0.259 / 0.283 / 4.012
12 020 11 972 1000 / 1000 0 % 0.278 / 0.301 / 0.559
13 020 12 972 1000 / 1000 0 % 0.298 / 0.319 / 0.634
13 478 13 430 1000 / 1000 0 % 0.307 / 0.325 / 0.654
13 520 13 472 1000 / 1000 0 % 0.306 / 0.326 / 0.541
14 020 13 972 1000 / 1000 0 % 0.317 / 0.336 / 0.579
15 020 14 972 1000 / 1000 0 % 0.324 / 0.348 / 1.079
16 000 15 952 1000 / 1000 0 % 0.345 / 0.366 / 0.671

Pi 4 end0 before / after the sweep: rx_errors 0 / 0, rx_crc_errors 0 / 0,
rx_length_errors 0 / 0, rx_align 0 / 0, tx_errors 0 / 0. RTT grows linearly with frame
size (about 12 µs per extra KB, i.e. wire time at 1 Gb/s) with no outliers at the old threshold.

Sizes 13 478 and 13 520 were chosen deliberately: 13 478 is where Snyder's Pi 4 started losing
packets at 1 Gb/s on the pre-fix build (report of 2026-09-14), and ~13 500 is where
@antonellocaroli saw his two Pi 4s become unstable. With bcm_phy_enable_jumbo() on both ends
neither threshold is visible any more; the BCM54213PE passes 16 000-byte frames at 1 Gb/s
despite the 10 KB figure on its product page.

2. Native DSD512 at 1000 Mb/s, MTU 16000 on both ends — real payload, both ends counted

Pings prove the PHY passes the frames; this proves the frames carry audio that plays. The Pi 4
has been streaming a native DSD512 album (DFF, 2 × 22.5792 MHz, ~46 Mb/s) to the CM4 since
08:27:34 — 8 tracks opened back to back in the renderer journal by 08:43, no gap, no dropout,
the DAC playing throughout. Diretta SDK profile negotiated with the CM4: cycle=2834us,
cycleSize=15992B, packets/cycle=1, reqMTU=16000 maxMTU=16000 — one Ethernet frame per cycle.

Counters read on both ends around the same 31 s window, mid-album (08:42:42 → 08:43:13):

Pi 4 (host, TX) CM4 (target, RX)
Frames 10 725 sent (346 per second) rx_pkts +10 955 (audio frames + Diretta control)
Bytes / average frame size 171 354 928 B / 15 977 B per frame rx_oversize +10 935 (one per jumbo frame, see MIB note)
Throughput 44.3 Mb/s (track change inside the window)
Errors tx_errors 0, tx_dropped unchanged rx_crc_errors 0, rx_frame_errors 0, rx_jabber 0, rx_length_errors 0
Software renderer: no underrun / xrun, SoC 61.8 °C, get_throttled=0x0 diretta_alsa_target active, end0 mtu 16000 / maxmtu 16347, 1000 Mb/s full

Every frame the Pi 4 put on the wire was received intact by the CM4: the 230-frame difference
is the target's own control traffic (rx_good_pkts +20 for the ≤1518-byte ones plus the
Diretta return channel). A shorter 10 s sample taken earlier at track steady state gave
362 frames per second, 16 003 bytes per frame, 46.4 Mb/s, end0 interrupts 374 per second
(all on CPU0), CPU 1.4–3.4 % per core, audio thread SCHED_FIFO 50 at 1.8 %.

The ping sweep of section 1 ran while this stream was playing; neither disturbed the other.

Note on the MIB statistics (same as on the CM4)

Every jumbo frame increments tx_oversize (host side) and rx_oversize (target side) by one
and no size bucket above 4096_9216 moves. On the Pi 4 after the sweep: rx_oversize +10 014
for the 10 000 echo replies, rx_4096_9216_oct +1 002 (only the 9 020-byte replies land in a
bucket), tx_oversize at 258 125 after a few minutes of DSD512. Cosmetic: CRC counters stay at
zero and the payload is intact.

Not covered here

100 Mb/s and 10 Mb/s on the Pi 4 side (the CM4 report already covers those speeds with the PHY
fix); can be run on request with the same script.

Thanks again — with the PHY fix, Pi 4 Model B is 16 K-clean at 1 Gb/s too.

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

Thanks for testing. I will have a look regarding the MIB counters, but am not really confident that this can be fixed as rx_oversize and rx_good_pkts come straight from the UMAC. The hardware counts everything above 1518 bytes as oversize and leaves it out of good_pkts, which is older than any jumbo support.

Functionality is not affected and the interface statistics are correct (rx_packets matches and no error counter), so I'd assume this is ok.

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

Do you now if the status block bug is in all GENET (non v1) versions? I only have v5 here to test.

Maybe the right way to think about it is that the original one descriptor per jumbo frame wasn't meant to work. It just so happened to work with the RSB disabled. So it is safe to assume this is the correct way to do things in all revisions of genet. The RTL also corroborates.

Thanks for taking this on. Good work so far!

Thanks for the confirmation. I will mention this in my cover letter, so clashiko does not have to ask 😄

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

@herisson-88 may I add a Tested-by: Name <email> to the upstream series for your testing?

@herisson-88

Copy link
Copy Markdown

If you want 🤗 but it is very small things...

@nbuchwitz

Copy link
Copy Markdown
Contributor Author

It is always good to see if someone else than the author has actually tested the patches. I would need the name and mail i should use for the tag

@herisson-88

Copy link
Copy Markdown

Pierre-Marin Leclercq pierremarinleclercq88@gmail.com

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants