- Replace build stub with full phase execution (prepare/build/check/install) via /bin/sh with variable interpolation, config file generation, and env handling (hard-set, soft/?=, append/+=) - Add += append operator for environment variables across DSL, parser, system config parser, build backend, and formatter - Add STYLEGUIDE.md documenting all codebase conventions - Replace EnvEntry bool soft with EnvMode enum (Set/Soft/Append) - Add Plus token type to lexer for += parsing
14 KiB
Style Guide
If you're reading this because you thought kappa's C++ looked different from what you're used to — good. That's the point.
This isn't a suggestion box. It's what the codebase looks like, and it's what your code will look like after you've rewritten it three times because the PR reviewer sent it back. Save yourself the rewrite. Read this first.
The Philosophy
We write C++ like it's the year 2026 and the committee finally shipped
something usable. No polyfills. No third-party libraries. No Boost. The
standard library is sufficient for a package manager. If you disagree, you
haven't read <format> closely enough.
Every line of kappa assumes the reader is competent. We don't explain what
std::string_view is. We don't annotate obvious control flow. Comments exist
to explain why, never what. If your code needs a comment to be
understood, the code is wrong.
Simplicity is a moral position. The scheduler is the hardest thing in this codebase, and it's 290 lines. If your feature adds more than that, you're building the wrong feature.
Naming
Structs, classes, enums
// PascalCase. Always.
struct BuildResult { };
enum class TokenType : std::uint8_t { };
// Enum values are PascalCase too. This isn't Java.
enum class EnvMode : std::uint8_t { Set, Soft, Append };
// Acronyms stay capitalized. B-Tree is BTree, not Btree.
// Two-letter acronyms stay capitalized. ID, not Id.
Type names state what the thing is, not what it's for. SchedResult, not
ResultForScheduler. InitPaths, not PathsForInitSystems.
Variables and functions
// snake_case. No Hungarian notation. No m_ prefix. No s_ prefix.
int pending_deps = 0;
void compute_depths(Scheduler& s);
std::string_view token_name(TokenType type);
Member variables and locals look identical. If you can't tell them apart, your functions are too long. Fix the function.
Files
src/service/openrc.cpp # snake_case, lowercase
include/kappa/service/types.hpp # .hpp for headers, .cpp for source
One public class per header is a myth invented by Java developers. Group
related declarations. types.hpp holds all enums and structs for a module.
If a module has one public struct and one public function, they go in the
same header.
Formatting
Indentation and braces
Four spaces. Attached braces (a K&R variant).
// ✓ yes — brace on the same line as the control structure
if (pid < 0) {
return -1;
}
// ✗ no — Allman/BSD braces on their own line
if (pid < 0)
{
return -1;
}
// ✗ no — missing braces on single-statement bodies
if (pid < 0) return -1;
Always braces. Even for single statements. The compiler doesn't care. The human reviewing your diff at 2 AM does. clang-tidy will flag bare bodies — apply the fix every time.
Line length
100 columns. Not 80 — we're not teletypes. Not 120 — if you need 120 characters to express a thought, your thought is too complicated. Split it.
Section separators
// --- Section description ---
// or
// ---------------------------------------------------------------------------
// Longer section description spanning the full runway
// ---------------------------------------------------------------------------
Thin lines (---) for sub-sections within a file. Thick lines
(-----------) for top-level section boundaries. The difference communicates
hierarchy without nesting.
Switch cases
switch (is) {
case InitSystem::Systemd:
return generate_systemd_service(spec);
case InitSystem::S6:
return generate_s6_service(spec);
case InitSystem::Unknown:
default:
return {};
}
Cases at the same indentation as the switch. No blocks around single-return
cases. Break or return in every non-fallthrough case. If you're falling
through intentionally, wrap it with [[fallthrough]]; on a line by itself.
Types
Use auto when the type is obvious, explicit when it isn't
// ✓ yes — type is obvious from initialization
auto scope = eval::make_default_scope();
auto& pkg = registry.at(name);
// ✓ yes — structured bindings, type is obvious
for (auto& [key, val] : features) { }
// ✓ yes — explicit where the type carries meaning
std::unique_lock lock(s.mtx); // not auto lock = ...
std::uint64_t h = 14695981039346656037ULL; // not auto h = ...
const is the default
Everything is const until proven mutable.
// ✓ yes
const auto& step = plan.steps[i];
for (const auto& entry : entries) { }
// ✗ no — mutable when it shouldn't be
auto& step = plan.steps[i];
View types over owning types in parameters
// ✓ yes
void set_root(std::string_view path);
bool is_supported(std::string_view name);
// ✗ no
void set_root(const std::string& path);
Return owning types. Accept views. The caller decides ownership. You decide what you need to read.
Strong enums only
// ✓ yes
enum class InitSystem : std::uint8_t { Systemd, OpenRC, S6 };
// ✗ no
enum InitSystem { INIT_SYSTEMD, INIT_OPENRC, INIT_S6 };
No unscoped enums. No ALL_CAPS enum values. No integer conversions without
explicit intent. If you need to serialize an enum to an integer, write a
to_string function. The enum's numeric value is an implementation detail,
not an interface.
Functions
One responsibility per function
If your function name contains the word "and", it does at least two things
and needs to be split. If the body doesn't fit on one screen, it does too
much. "One screen" means approximately 30 lines. The scheduler's
compute_depths is 30 lines. So is the resolver's resolve. They're at the
upper bound. If yours is longer, you're doing something wrong.
Error handling: return, don't throw
// ✓ yes
struct FetchResult {
std::filesystem::path work_dir;
std::string error;
bool ok() const { return error.empty(); }
};
FetchResult fetch(const PackageDef& pkg);
// ✗ no
void fetch(const PackageDef& pkg); // throws on error
Exceptions are for unrecoverable programmer errors — out-of-memory, null
dereference, violated invariants. They belong in constructors and in the
parser (where ParseError is the only way to unwind back to diagnostics).
Every operational failure — network down, disk full, configure script failed,
hash mismatch — is a return value. A struct with bool ok and
std::string error. Check the ok field, read the error string, don't
catch exceptions for normal operation.
Return early, return often
// ✓ yes
if (to.empty()) {
return {false, "destination is empty"};
}
// ... main logic ...
// ✗ no
if (!to.empty()) {
// ... 40 lines of nesting ...
} else {
return {false, "destination is empty"};
}
Guard clauses at the top. Happy path straight down the left margin. If your code has three levels of nesting, you missed an early return opportunity.
Static helpers over lambdas
If a helper is more than 5 lines, extract it to a file-static function above
the public API. Named functions are greppable. Named functions show up in
stack traces. Lambdas don't. The one exception is a run_phase lambda in
build() — it captures local state that would require a 5-parameter helper
and it's clearly a one-off control flow wrapper, not a reusable abstraction.
Namespaces
namespace kappa::module {
// Everything goes here.
} // namespace kappa::module
C++17 nested namespace syntax. Closing brace gets a comment with the namespace name. These comments survive diffs, refactors, and editors that collapse braces. They cost one line and save ten minutes of scrolling up to figure out which brace closes what.
No using namespace at file scope
// ✓ yes — inside a function
namespace fs = std::filesystem;
// ✗ no — at file scope
using namespace std;
Namespace aliases are acceptable inside functions — namespace fs = std::filesystem; is fine when the file does a lot of path manipulation. But
at file scope? No. You're not writing using namespace std; at the top of a
header and you're not doing the subtler version of the same sin.
Headers
#pragma once
#include "kappa/resolve/plan.hpp"
#include "kappa/dsl/ast.hpp"
#include <string>
#include <vector>
namespace kappa::build {
struct BuildResult { };
BuildResult build(const resolve::BuildStep& step,
const std::string& work_dir,
int jobs);
} // namespace kappa::build
#pragma once at the top. No include guards. This is 2026.
Project headers first, in quotes. System headers second, in angle brackets. Blank line between the two groups. Alphabetical within each group.
Headers include only what they need to compile. If build.hpp uses
resolve::BuildStep by reference, it includes resolve/plan.hpp. It does
not forward-declare BuildStep — we don't forward-declare across module
boundaries. The include is the contract: "this module depends on that one."
Headers never contain implementation. No inline functions. No
template definitions in headers (we don't use templates). The one exception
is parse_util.hpp, which defines ParseError inline because it's a thin
exception wrapper and splitting it would be ceremony for ceremony's sake.
One exception per codebase is a pattern. Two is a problem.
Modules
Every module follows this structure:
include/kappa/{module}/
├── types.hpp # enums, structs, parse/validate declarations
├── {feature}.hpp # public function declarations
src/{module}/
├── types.cpp # parse/validate/describe implementations
├── backend_a.cpp # per-variant generation (if applicable)
├── backend_b.cpp
└── install.cpp # dispatch + orchestration (if applicable)
If a module doesn't need types.hpp (single struct, single function), both
go in {feature}.hpp. If a module has no backends, skip them. But don't
invent a third pattern. service/ and boot/ are the templates. Copy them.
Strings and formatting
// ✓ yes
auto msg = std::format("building {} (depth={})", name, depth);
result.error = std::format("command exited with code {}: {}", rc, cmd);
// ✗ no — ostringstream for trivial concatenation
std::ostringstream oss;
oss << "building " << name << " (depth=" << depth << ")";
// ✓ yes — ostringstream for incremental construction
std::ostringstream out;
out << "[Unit]\n";
out << std::format("Description={}\n", desc);
std::format for one-shot strings. std::ostringstream for building up
output incrementally (service files, bootloader configs, formatter output).
String concatenation with + is acceptable for two or three pieces.
Anything more goes through std::format.
String views for parameters
// ✓ yes
InitSystem parse_init_system(std::string_view name);
void print_error(std::ostream& os, std::string_view source,
SourceLocation loc, std::string_view message);
// ✗ no
InitSystem parse_init_system(const std::string& name);
Views everywhere, except when you need to store the string.
The DSL
The .kap DSL grammar is the contract. You can extend it. You cannot break
existing configs. Every new token type requires:
- An entry in
TokenType - A case in
token_name() - Parsing logic in the appropriate parser
- A formatting case in
format.cpp - At least one test in
test.shthat exercises the new syntax
If you're adding a keyword, think twice. The lexer already has 29 token types. Every new one increases parse time and mental overhead. Can this be expressed with the existing grammar? If yes, don't add a keyword.
Thread safety
The scheduler is multithreaded. If you touch shared state, you own the lock.
{
std::unique_lock lock(s.mtx);
s.waiting.erase(idx);
}
// lock released here — no shared state access beyond this point
Use scoped locks. Never lock/unlock manually. Never hold a lock across a
condition variable wait without understanding why. If you think you need
memory_order_release, you probably need memory_order_acq_rel and you
should document why in a three-line comment above the operation.
If a data structure is touched by multiple threads, its access pattern must be documented at the declaration site, not in a PR description. "This is only written under the lock, read atomically elsewhere" goes in the header.
What clang-tidy enforces
We run with -Wall -Wextra -Wpedantic and a .clang-tidy config. Zero
warnings. Not "zero warnings except for that one file." Zero.
The following are non-negotiable:
- Every
if/for/whilebody has braces autovariables that are never modified areconst auto- Variables are initialized at declaration
- No unused includes
- No redundant declarations
If clang-tidy suggests a fix and you disagree, you're wrong. Apply the fix.
The only acceptable override is // NOLINT with a justification comment —
and if you write that more than twice in a file, the reviewer will ask you to
rethink your design.
What we reject
-
Comments that narrate the code.
// Increment counterabovei++is an insult. Delete it. -
Dead code. No commented-out blocks. No
#if 0. If it's not used, it doesn't exist. Git remembers. -
Premature abstraction. Three identical lines do not need a function. Ten do. The threshold is somewhere in between and you should err on the side of duplication.
-
C heritage.
printf,malloc,NULL, rawchar*strings,#defineconstants. The 1970s called. Don't answer. -
Over-engineering. The build backend doesn't need a plugin architecture. The lexer doesn't need a state machine framework. Solve the problem in front of you, not the one you imagine someone might have in three years.
-
Cleverness. If your solution makes you feel smart, it's wrong. The best code is the code you forget about because it never breaks.
Kappa does one thing: build your system from source, init-agnostically. Everything in this style guide exists to keep that codebase small, fast, and comprehensible. If a rule conflicts with that goal, the goal wins — but you'd better have a good reason, and you'd better write it down.