184 lines
9.0 KiB
Markdown
184 lines
9.0 KiB
Markdown
# fastwc
|
|
|
|
**The word counter GNU wishes it could be.**
|
|
|
|
GNU coreutils `wc` is the Apple of the FOSS world. Forty years of
|
|
accumulated bureaucracy wrapped in a binary. Translation teams. gettext.
|
|
`--help` output in fourteen languages. An autotools contraption the size
|
|
of a small city, all so you can count newlines. And when it can't keep
|
|
up, it doesn't get faster — it gets *more dependencies*.
|
|
|
|
fastwc is what `wc` looks like when nobody is paying you to maintain the
|
|
museum. One file. One purpose. No translators. No gnulib. No AVX-512
|
|
kernels hand-tuned by people whose entire job is compensating for the
|
|
bloat around them — just our own: AVX-512, AVX-2, and SSE2 intrinsics
|
|
with runtime dispatch, and a scalar SWAR fallback. Just counting,
|
|
correctly, at full speed.
|
|
|
|
## The scoreboard
|
|
|
|
The benchmark suite in `benchmarks/` races fastwc against GNU `wc`,
|
|
busybox `wc`, and toybox `wc` — fail-fast. The moment we are slower, or
|
|
disagree on a single count, it writes a shame report and exits
|
|
non-zero. These are the facts. Earlier numbers were measured on a
|
|
workstation with more cores than taste; these are from an ordinary
|
|
laptop APU, which is where the rest of you live:
|
|
|
|
| Suite | Result |
|
|
|-------|--------|
|
|
| words, tiny (1-10k lines) | **6/6 vs coreutils, busybox, and toybox.** Correct on every one — but these races finish in the low hundreds of microseconds and prove *right*, not *fast*. See the note on startup below. |
|
|
| words, 1M lines (11 MB) | **1.6 ms vs coreutils 7.9 ms (4.9x).** Busybox 10.0x (16.2 ms), toybox 9.9x (15.9 ms). |
|
|
| words, 10M lines (110 MB) | **5.1 ms vs coreutils 66 ms (13.1x).** Busybox 27.0x (160 ms), toybox 15.9x (157 ms). |
|
|
| lines, 1M (11 MB) | **1.6 ms vs busybox 16.9 ms (10.4x)**, toybox 13.3x (17.6 ms). vs coreutils this is startup, not throughput — see below. |
|
|
| lines, 10M (110 MB) | **4.6 ms vs coreutils 5.8 ms (1.27x).** GNU throws every thread at a mapped file and the scoreboard doesn't move. Busybox 20.1x (160 ms), toybox 31.3x (154 ms). |
|
|
| lines, 100M monster | **29.2 ms vs coreutils 47.5 ms (1.6x).** The bigger the file, the wider the gap. |
|
|
| lines, 1B solo | **266 ms.** 11 GB in one pass at 41.4 GB/s — 3.76 billion lines per second. We are the reference now. |
|
|
| stdin lines, 10M | **4.6 ms vs coreutils 5.9 ms (1.28x).** Busybox 16.6x (161 ms), toybox 16.8x (154 ms). |
|
|
| stdin words, 1M | **1.5 ms vs coreutils 7.1 ms (4.7x).** GNU reads stdin like it's 1985. Busybox 12.7x, toybox 10.4x. |
|
|
|
|
Measured on an AMD Ryzen AI 7 PRO 350 — 8 cores / 16 threads, boost up
|
|
to 5.09 GHz, 384 KiB L1d / 256 KiB L1i / 8 MiB L2 / 16 MiB L3 — with
|
|
64 GiB of DDR5-4800 dual-channel (2x 32 GiB SODIMM). Opponents: GNU
|
|
coreutils 9.11, busybox 1.36.1, toybox 0.8.13. Page-cache-warm files,
|
|
three interleaved runs per case, minimum kept, both sides timed at
|
|
microsecond resolution by a C timer (`benchmarks/tools/timeit`) — no
|
|
`date`-fork rounding. Cases where the reference itself finishes in
|
|
under 5 ms are startup, not throughput: fastwc must still match the
|
|
count, and the case is reported, but it is excluded from the averages.
|
|
Raced cases allow a 2% dead-heat margin so a genuine tie can't flake
|
|
on scheduler jitter. All of it, exactly as `benchmarks/` prescribes.
|
|
|
|
The moment fastwc is slower than any of them, this project has failed
|
|
and you should say so loudly in an issue. The benchmark is the
|
|
contract. The how and why of the speed, with receipts, lives in
|
|
[docs/PERFORMANCE.md](docs/PERFORMANCE.md).
|
|
|
|
Toybox joined the suite for the same reason busybox is still there:
|
|
shits and giggles. It stopped giggling around the first average —
|
|
~15x — and it has never won a single case; counting 10M lines takes
|
|
it 16 times as long as us, and busybox 27. GNU coreutils is the only
|
|
opponent that keeps score, and even its threaded counter only manages
|
|
a dead heat on mid-size files — never a win, and the moment the file
|
|
stops fitting in a polite buffer, the dead heat stops being polite.
|
|
|
|
A note on startup, in the interest of honesty: on a one-line file the
|
|
whole race happens in the low hundreds of microseconds — fastwc
|
|
~0.25 ms, GNU ~0.33 ms, toybox ~0.37 ms, busybox ~0.14 ms. Microseconds
|
|
either way. Nobody will ever notice a difference that small, and it
|
|
does not matter in the bigger picture: the tiny cases in the table are
|
|
here to prove fastwc is never *wrong*, not to brag about a head start
|
|
that evaporates the moment the page cache warms up. That is why the
|
|
suite now times both sides at microsecond resolution and files anything
|
|
the reference finishes in under 5 ms under "startup-bound": correct,
|
|
reported, and excluded from the averages — because nobody should be
|
|
racing startup, least of all a word counter. The cases that matter are
|
|
the ones where counting takes longer than starting — and those are the
|
|
ones in this table. Every run ends with the average speedup against
|
|
each oracle — coreutils ~4.5x, busybox ~13-15x, toybox ~15-18x — so
|
|
the cruelty is quantified.
|
|
|
|
## Why
|
|
|
|
- **GNU wc is a dependency museum.** Its build needs gettext, gnulib,
|
|
and a translator for every language on Earth. fastwc needs `cc`.
|
|
- **GNU wc is slow where it should be fast.** Counting bytes is not
|
|
supposed to be an architectural achievement. Counting 11 GB of them
|
|
in one pass, solo, while GNU's AVX-512 assembly still needs a buffer
|
|
to copy into — that is.
|
|
- **GNU wc counts like it's 1985** — because it is. We count like it's
|
|
now: regular files are mapped and counted in parallel across cores,
|
|
with SIMD kernels (AVX-512, AVX-2, SSE2) dispatched at runtime —
|
|
zero function calls in the hot path.
|
|
|
|
## What it does
|
|
|
|
```
|
|
fastwc [OPTION]... [FILE]...
|
|
fastwc [OPTION]... --files0-from=F
|
|
```
|
|
|
|
- `-l` lines, `-w` words, `-c` bytes, `-m` characters (multibyte),
|
|
`-L` maximum line length (display width, tab stops every 8)
|
|
- `--files0-from=F` NUL-terminated file lists from F (or stdin when F is `-`)
|
|
- `--total=auto|always|only|never` control over the total row
|
|
- `--debug` report which line-counting acceleration is in use
|
|
- stdin, `-`, multiple files, `total` rows, GNU-compatible counts
|
|
- A CLI surface modeled on GNU coreutils `wc` 9.11: same options, same
|
|
column alignment rule (counts are right-justified to the widest sum
|
|
of regular file sizes), same multibyte word semantics, same error
|
|
text. Swap it in and scripts keep working.
|
|
- POSIXLY_CORRECT is honored: the no-break space stops being a word
|
|
separator, exactly like GNU wc.
|
|
- no `--help` in fourteen languages. One `--help`, in English, the
|
|
language of people who ship software
|
|
|
|
## Build
|
|
|
|
Requires a C compiler and autotools. That's it. No gettext. No gnulib.
|
|
No translators.
|
|
|
|
```sh
|
|
./autogen.sh # autoreconf -fi && ./configure
|
|
make
|
|
make release # installs the release binary to bin/release/fastwc
|
|
```
|
|
|
|
## Benchmark
|
|
|
|
```sh
|
|
make bench # build release + run every suite
|
|
./benchmarks/test-all.sh # all suites: words, lines, stdin
|
|
```
|
|
|
|
The suites live under `benchmarks/files/{lines,words}` (file input) and
|
|
`benchmarks/stdin/piping` (standard input). Every suite races fastwc
|
|
against GNU coreutils wc and, when it is installed, busybox wc and
|
|
toybox wc. The benchmark locates the real oracles by their `--version`
|
|
answer — a `wc` that answers as fastwc is a symlink somebody made, and
|
|
is skipped rather than raced against itself.
|
|
Both sides are timed at microsecond resolution by `tools/timeit`, a C
|
|
timer built on the spot (`test-all.sh` compiles it next to `genfile`)
|
|
that forks the command, discards its output, and reads the clock after
|
|
it exits — no `date`-fork rounding. The suites interleave runs so both
|
|
commands see identical cache warmth and keep the minimum of three.
|
|
A case where the reference itself finishes in under 5 ms is startup,
|
|
not throughput: fastwc must still match the count, but the case is
|
|
reported as `startup-bound` and excluded from the averages. Raced
|
|
cases allow a 2% dead-heat margin so a genuine tie can't flake — but a
|
|
real loss writes a shame report, fails the suite, and exits non-zero.
|
|
The words suite grows to 1M and 10M lines; the lines suite carries the
|
|
monsters: 100M lines raced against coreutils, and 1B lines timed solo
|
|
— there is no reference for that one; we are the reference now. When
|
|
the run ends, the average speedup of fastwc against each oracle is
|
|
printed.
|
|
|
|
## Development
|
|
|
|
The editor setup is one command:
|
|
|
|
```sh
|
|
make compile-commands # compile_commands.json for clangd
|
|
```
|
|
|
|
clangd reads `.clangd`, `.clang-tidy`, and `.clang-format` — the style
|
|
guide, enforced by robots. We use `bear` when it's installed; the
|
|
fallback hand-rolls the single entry from the Makefile, because one
|
|
source file doesn't need a database.
|
|
|
|
- `make format` — make the code confess to the style guide
|
|
- `make format-check` — verify without touching
|
|
- `make lint` — clang-tidy, static analysis included
|
|
|
|
`.editorconfig` and `.gitattributes` keep every editor honest. Your
|
|
editor has opinions. So do we. Ours are in the repo.
|
|
|
|
## License
|
|
|
|
MIT. Do whatever you want. We're not GNU, we won't sue you — we'll just
|
|
be faster.
|
|
|
|
## Contributing
|
|
|
|
See [CONTRIBUTING.md](CONTRIBUTING.md) — the rules are the deal.
|
|
See [STYLEGUIDE.md](STYLEGUIDE.md) — the style is the law.
|