Files

11 KiB

fastwc

The word counter GNU wishes it could be.

GNU coreutils wc is the Apple of the FOSS world. Forty years of accumulated bureaucracy wrapped in a binary. Translation teams. gettext. --help output in fourteen languages. An autotools contraption the size of a small city, all so you can count newlines. And when it can't keep up, it doesn't get faster — it gets more dependencies.

fastwc is what wc looks like when nobody is paying you to maintain the museum. One file. One purpose. No translators. No gnulib. No AVX-512 kernels hand-tuned by people whose entire job is compensating for the bloat around them — just our own: AVX-512, AVX-2, and SSE2 intrinsics with runtime dispatch, and a scalar SWAR fallback. Just counting, correctly, at full speed.

The scoreboard

The benchmark suite in benchmarks/ races fastwc against GNU wc, busybox wc, and toybox wc — fail-fast. The moment we are slower, or disagree on a single count, it writes a shame report and exits non-zero. These are the facts. Earlier numbers were measured on a workstation with more cores than taste; these are from an ordinary laptop APU, which is where the rest of you live:

Suite Result
words, tiny (1-10k lines) 6/6 vs coreutils, busybox, and toybox. Correct on every one — but these races finish in the low hundreds of microseconds and prove right, not fast. See the note on startup below.
words, 1M lines (11 MB) 1.6 ms vs coreutils 7.9 ms (4.9x). Busybox 10.0x (16.2 ms), toybox 9.9x (15.9 ms).
words, 10M lines (110 MB) 5.1 ms vs coreutils 66 ms (13.1x). Busybox 27.0x (160 ms), toybox 15.9x (157 ms).
lines, 1M (11 MB) 1.6 ms vs busybox 16.9 ms (10.4x), toybox 13.3x (17.6 ms). vs coreutils this is startup, not throughput — see below.
lines, 10M (110 MB) 4.6 ms vs coreutils 5.8 ms (1.27x). GNU throws every thread at a mapped file and the scoreboard doesn't move. Busybox 20.1x (160 ms), toybox 31.3x (154 ms).
lines, 100M monster 29.2 ms vs coreutils 47.5 ms (1.6x). The bigger the file, the wider the gap.
lines, 1B solo 266 ms. 11 GB in one pass at 41.4 GB/s — 3.76 billion lines per second. We are the reference now.
stdin lines, 10M 4.6 ms vs coreutils 5.9 ms (1.28x). Busybox 16.6x (161 ms), toybox 16.8x (154 ms).
stdin words, 1M 1.5 ms vs coreutils 7.1 ms (4.7x). GNU reads stdin like it's 1985. Busybox 12.7x, toybox 10.4x.

Measured on an AMD Ryzen AI 7 PRO 350 — 8 cores / 16 threads, boost up to 5.09 GHz, 384 KiB L1d / 256 KiB L1i / 8 MiB L2 / 16 MiB L3 — with 64 GiB of DDR5-4800 dual-channel (2x 32 GiB SODIMM). Opponents: GNU coreutils 9.11, busybox 1.36.1, toybox 0.8.13. Page-cache-warm files, three interleaved runs per case, minimum kept, both sides timed at microsecond resolution by a C timer (benchmarks/tools/timeit) — no date-fork rounding. Cases where the reference itself finishes in under 5 ms are startup, not throughput: fastwc must still match the count, and the case is reported, but it is excluded from the averages. Raced cases allow a 2% dead-heat margin so a genuine tie can't flake on scheduler jitter. All of it, exactly as benchmarks/ prescribes.

And every number above was measured in the C locale — the setting that flatters the opponents most. Under en_US.UTF-8, GNU wc stops counting bytes and starts decoding them, one mbrtowc at a time, even when the file is pure ASCII and decoding changes nothing. We used to make that exact mistake: a UTF-8 locale silently swapped our SIMD kernels for the same decoder, and the 11 MB words race flipped from a 4.9x win to a GNU win. The kernels now probe for non-ASCII bytes while they count — a vector move-mask per load, free when unused — so ASCII files never see the decoder. Same 11 MB words file, the locale you actually run: 1.4 ms vs GNU 10.4 ms (7.6x). GNU still pays that 10.4 ms for bytes that were never multibyte; we pay for the decoder only when a file genuinely needs it. Receipts in docs/PERFORMANCE.md.

The moment fastwc is slower than any of them, this project has failed and you should say so loudly in an issue. The benchmark is the contract. The how and why of the speed, with receipts, lives in docs/PERFORMANCE.md.

Toybox joined the suite for the same reason busybox is still there: shits and giggles. It stopped giggling around the first average — ~15x — and it has never won a single case; counting 10M lines takes it 16 times as long as us, and busybox 27. GNU coreutils is the only opponent that keeps score, and even its threaded counter only manages a dead heat on mid-size files — never a win, and the moment the file stops fitting in a polite buffer, the dead heat stops being polite.

A note on startup, in the interest of honesty — and of gloating: on a one-line file the whole race happens in the low hundreds of microseconds, and fastwc now wins it outright. The default build is static musl, so there is no dynamic loader to pay: min-of-400 on a 12-byte file puts fastwc -l at 78µs against busybox's 83µs, GNU's 253µs, and toybox's 254µs. Busybox's one structural advantage — a loader it never had to start — is no longer an advantage; we don't start one either. None of this matters in the bigger picture, and we will not pretend otherwise: nobody will ever notice a difference that small, and the tiny cases in the table are here to prove fastwc is never wrong, not to brag about a head start that evaporates the moment the page cache warms up. That is why the suite now times both sides at microsecond resolution and files anything the reference finishes in under 5 ms under "startup-bound": correct, reported, and excluded from the averages — because nobody should be racing startup, least of all a word counter. The cases that matter are the ones where counting takes longer than starting — and those are the ones in this table. Every run ends with the average speedup against each oracle — coreutils ~4.5x, busybox ~13-15x, toybox ~15-18x — so the cruelty is quantified.

Why

  • GNU wc is a dependency museum. Its build needs gettext, gnulib, and a translator for every language on Earth. fastwc needs cc.
  • GNU wc is slow where it should be fast. Counting bytes is not supposed to be an architectural achievement. Counting 11 GB of them in one pass, solo, while GNU's AVX-512 assembly still needs a buffer to copy into — that is.
  • GNU wc counts like it's 1985 — because it is. We count like it's now: regular files are mapped and counted in parallel across cores, with SIMD kernels (AVX-512, AVX-2, SSE2) dispatched at runtime — zero function calls in the hot path.
  • GNU wc slows down in the locale you actually run. Under a UTF-8 locale it decodes every byte it counts — pure ASCII included, which decoding cannot change — so the 11 MB words file that takes it 7.9 ms in the C locale takes 10.4 ms there, against our 1.4 ms. fastwc's kernels probe for non-ASCII bytes as they count and only decode files that need it. The locale that taxes GNU is free for us.

What it does

fastwc [OPTION]... [FILE]...
fastwc [OPTION]... --files0-from=F
  • -l lines, -w words, -c bytes, -m characters (multibyte), -L maximum line length (display width, tab stops every 8)
  • --files0-from=F NUL-terminated file lists from F (or stdin when F is -)
  • --total=auto|always|only|never control over the total row
  • --debug report which line-counting acceleration is in use
  • stdin, -, multiple files, total rows, GNU-compatible counts
  • A CLI surface modeled on GNU coreutils wc 9.11: same options, same column alignment rule (counts are right-justified to the widest sum of regular file sizes), same multibyte word semantics, same error text. Swap it in and scripts keep working.
  • POSIXLY_CORRECT is honored: the no-break space stops being a word separator, exactly like GNU wc.
  • no --help in fourteen languages. One --help, in English, the language of people who ship software

Build

Requires a C compiler and autotools. That's it. No gettext. No gnulib. No translators.

The default build links statically against musl (via musl-gcc), which is why fastwc now wins the startup cases above outright — there is no dynamic loader to pay, and the ~80µs exec floor is the same one busybox pays. If musl-gcc isn't installed the configure script warns and falls back to the system compiler; the glibc build is one flag away:

./autogen.sh            # autoreconf -fi && ./configure (static musl)
make
make release            # installs the release binary to bin/release/fastwc

./configure --enable-glibc && make   # or: dynamic glibc build

Benchmark

make bench                          # build release + run every suite
./benchmarks/test-all.sh            # all suites: words, lines, stdin

The suites live under benchmarks/files/{lines,words} (file input) and benchmarks/stdin/piping (standard input). Every suite races fastwc against GNU coreutils wc and, when it is installed, busybox wc and toybox wc. The benchmark locates the real oracles by their --version answer — a wc that answers as fastwc is a symlink somebody made, and is skipped rather than raced against itself. Both sides are timed at microsecond resolution by tools/timeit, a C timer built on the spot (test-all.sh compiles it next to genfile) that forks the command, discards its output, and reads the clock after it exits — no date-fork rounding. The suites interleave runs so both commands see identical cache warmth and keep the minimum of three. A case where the reference itself finishes in under 5 ms is startup, not throughput: fastwc must still match the count, but the case is reported as startup-bound and excluded from the averages. Raced cases allow a 2% dead-heat margin so a genuine tie can't flake — but a real loss writes a shame report, fails the suite, and exits non-zero. The words suite grows to 1M and 10M lines; the lines suite carries the monsters: 100M lines raced against coreutils, and 1B lines timed solo — there is no reference for that one; we are the reference now. When the run ends, the average speedup of fastwc against each oracle is printed.

Development

The editor setup is one command:

make compile-commands   # compile_commands.json for clangd

clangd reads .clangd, .clang-tidy, and .clang-format — the style guide, enforced by robots. We use bear when it's installed; the fallback hand-rolls the single entry from the Makefile, because one source file doesn't need a database.

  • make format — make the code confess to the style guide
  • make format-check — verify without touching
  • make lint — clang-tidy, static analysis included

.editorconfig and .gitattributes keep every editor honest. Your editor has opinions. So do we. Ours are in the repo.

License

MIT. Do whatever you want. We're not GNU, we won't sue you — we'll just be faster.

Contributing

See CONTRIBUTING.md — the rules are the deal. See STYLEGUIDE.md — the style is the law.