huntedbytheirs ff465e981b perf: make gnu wc the slowest thing in the benchmark
SIMD kernels (AVX-512/AVX-2/SSE2, runtime-dispatched) counting a
single pass over mmap'd files, split across cores past 8 MiB, with
exact GNU oracle parity (NBSP included, glibc's decoder fixed, the
whole -m path mirrored so counts agree at every boundary).

Result: 1ms vs 2ms at 1M lines, 8-9ms vs 22-24ms at 10M. Forty years
of dependencies, hand-tuned AVX-512 assembly, a translation team per
language — and gnu wc still needs a buffer to copy into before it
can count. We mapped the file and just counted. The benchmark suite
no longer has a losing row; the shame report file is going to rust.
2026-08-29 15:32:08 -04:00
2026-08-29 13:48:15 -04:00
2026-08-29 13:57:59 -04:00
2026-08-29 13:57:59 -04:00
2026-08-29 13:57:59 -04:00
2026-08-29 13:57:59 -04:00
2026-08-29 13:57:59 -04:00
2026-08-29 13:57:59 -04:00
2026-08-29 13:57:59 -04:00
2026-08-29 13:48:15 -04:00
2026-08-29 13:48:15 -04:00
2026-08-29 12:36:12 -04:00

fastwc

The word counter GNU wishes it could be.

GNU coreutils wc is the Apple of the FOSS world. Forty years of accumulated bureaucracy wrapped in a binary. Translation teams. gettext. --help output in fourteen languages. An autotools contraption the size of a small city, all so you can count newlines. And when it can't keep up, it doesn't get faster — it gets more dependencies.

fastwc is what wc looks like when nobody is paying you to maintain the museum. One file. One purpose. No translators. No gnulib. No AVX-512 kernels hand-tuned by people whose entire job is compensating for the bloat around them — just our own: AVX-512, AVX-2, and SSE2 intrinsics with runtime dispatch, and a scalar SWAR fallback. Just counting, correctly, at full speed.

The scoreboard

The benchmark suite in benchmarks/ races fastwc against GNU wc (and busybox, if you keep such things installed) — fail-fast. The moment we are slower, or disagree on a single count, it writes a shame report and exits non-zero. These are the facts:

Suite Result
words (6 cases) 6/6 wins. Never slower, never wrong.
lines (up to 100k lines) Wins. GNU never sees us coming.
lines (1M lines) Win: 1ms vs 2ms. GNU's AVX-512 assist can't beat a mapped file.
lines (10M lines) Win: 8-9ms vs 22-24ms (~2.5x). GNU's lead never survives contact with the buffer.
busybox lines (10M) Win: 8-9ms vs ~165ms (~18x). If you must.

The moment fastwc is slower than GNU wc, this project has failed and you should say so loudly in an issue. The benchmark is the contract. The how and why of the speed, with receipts, lives in docs/PERFORMANCE.md.

Why

  • GNU wc is a dependency museum. Its build needs gettext, gnulib, and a translator for every language on Earth. fastwc needs cc.
  • GNU wc is slow where it should be fast. Counting bytes is not supposed to be an architectural achievement.
  • GNU wc counts like it's 1985 — because it is. We count like it's now: regular files are mapped and counted in parallel across cores, with SIMD kernels (AVX-512, AVX-2, SSE2) dispatched at runtime — zero function calls in the hot path.

What it does

fastwc [-lwc] [-m] [file...]
  • -l lines, -w words, -c bytes, -m characters (multibyte)
  • stdin, -, multiple files, total rows, GNU-compatible counts
  • no --help in fourteen languages. One --help, in English, the language of people who ship software

Build

Requires a C compiler and autotools. That's it. No gettext. No gnulib. No translators.

./autogen.sh     # autoreconf -fi && ./configure
make
make release     # installs the release binary to bin/release/fastwc

Benchmark

make bench                          # build release + run every suite
./benchmarks/bench-coreutils.sh     # the real fight
./benchmarks/bench-busybox.sh       # if you must

The suites interleave runs so both commands see identical cache warmth, keep the minimum, and fail the moment fastwc loses a single case. GNU wc is used as an oracle the same way you'd use a broken clock: occasionally it's right, and it's the only one around.

Development

The editor setup is one command:

make compile-commands   # compile_commands.json for clangd

clangd reads .clangd, .clang-tidy, and .clang-format — the style guide, enforced by robots. We use bear when it's installed; the fallback hand-rolls the single entry from the Makefile, because one source file doesn't need a database.

  • make format — make the code confess to the style guide
  • make format-check — verify without touching
  • make lint — clang-tidy, static analysis included

.editorconfig and .gitattributes keep every editor honest. Your editor has opinions. So do we. Ours are in the repo.

License

MIT. Do whatever you want. We're not GNU, we won't sue you — we'll just be faster.

Contributing

See CONTRIBUTING.md — the rules are the deal. See STYLEGUIDE.md — the style is the law.

S
Description
A extremely fast wc replacement.
Readme MIT
278 KiB
Languages
C 69.6%
Shell 26.8%
M4 1.5%
Makefile 1.3%
Nix 0.8%