Skip to content
andrew.dunn.dev

Building Bootc Images from Scratch

The official Fedora bootc image ships 523 packages. NetworkManager, chrony, firewalld, SSSD, sos, man-db, irqbalance. If you run a container host where every application is a Podman container, most of those packages serve no purpose. They are attack surface, update churn, and image weight.

Our immutable home servers layer on top of fedora-bootc:42. The first thing the Containerfile does is remove packages the upstream put in: dnf remove chrony firewalld plymouth. This is fragile. Upstream adds a new dependency, it sneaks back in. You are fighting the base image instead of controlling it.

The question: can you build a bootc image from scratch using only dnf --installroot and a standard Containerfile? No rpm-ostree compose. No treefile. No inheritance from the upstream image. Just packages you chose, transforms you understand, and a FROM scratch final stage.

The answer is yes. We built one, booted it, hardened it through 11 categories of validation, layered ZFS kernel modules on top without modification, and tested bootc upgrade and rollback against a live registry. This post documents the approach, what broke, what we learned about the upstream package tiers, and the final results.

Why not rpm-ostree compose?

The upstream Fedora and CentOS bootc images are built with rpm-ostree compose rootfs using treefile manifests. The treefile format gives you passwd/group management, postprocess scripts, recommends: false globally, and package removal. It is what the distribution maintainers use.

We wanted dnf --installroot for three reasons. First, it is a standard Fedora tool that produces a plain rootfs without ostree-specific tooling in the build path. Second, it keeps the Containerfile readable: every package and every transform is a visible RUN command, not buried in a YAML manifest processed by a specialized tool. Third, it tests whether the bootc ecosystem has decoupled enough from rpm-ostree that you can build images without it. If the answer is yes, the barrier to entry for custom bootc images drops significantly.

The upstream work

Before building the image, we submitted upstream tooling to make this easier for everyone. bootc container finalize-rootfs applies the ostree filesystem layout transforms (toplevel symlinks, /var cleanup, rpmdb relocation, config injection) to a dnf --installroot output. bootc container post-chroot-cleanup handles the artifacts left by chroot operations like dracut and bootupd.

For this build, we did not use those subcommands. They are not merged yet. Instead, we replicated the transforms as shell commands in the Containerfile, using the PR source as documentation for exactly what each transform does. When the PR merges, the Containerfile simplifies to two commands.

The package list

255 packages in three tiers, informed by the upstream Fedora bootc base-images tier structure:

Tier 0 — Bootable OS. The minimum to boot and manage the system: kernel, systemd, systemd-pam, bootc, bootupd, ostree, nss-altfiles, selinux-policy-targeted, container-selinux, coreutils, dnf, xfsprogs, e2fsprogs, dosfstools, dracut, tpm2-tools, grub2, grub2-efi-x64, efibootmgr, shim, microcode_ctl, linux-firmware, bubblewrap.

Tier 1 — Pure systemd base. Our opinionated choices replacing the upstream defaults, plus essential admin tools: systemd-boot-unsigned, systemd-networkd (replaces NetworkManager), systemd-resolved, nftables (replaces firewalld), openssh-server, openssh-clients, sudo, polkit, iproute, less, tar, hostname, vim-minimal, tzdata, attr, jq, bash-completion, audit.

Tier 2 — Container host. Every application is a container, so the container runtime is part of the base: podman, buildah, skopeo, crun, fuse-overlayfs.

TIER 2~5 pkgsContainer hostpodman, buildah, skopeo, crun255 PACKAGES INSTALLEDfedora-bootc:42 ships 523TIER 1~70 pkgsPure systemd basenetworkd over NM, nftables over firewalldsystemd-networkd, systemd-resolved, nftables, openssh, sudo, iproute, jq, auditTIER 0~180 pkgsBootable OSthe minimum to boot, deploy, and managekernel, systemd, bootc, bootupd, ostree, SELinux, dracut, grub2, shim, dnf, bubblewrap

Tier 0 carries the weight: booting and managing the machine costs roughly 180 of the 255 packages. The systemd opinions and the container runtime ride on top for another 75, and the whole base still lands under half of upstream’s 523.

What we dropped compared to fedora-bootc:42: NetworkManager, chrony, firewalld, plymouth, SSSD, sos, man-db, nano, irqbalance, WALinuxAgent-udev, nfs-utils, iptables-services, zram-generator, fwupd, and roughly 270 other packages.

The build

The Containerfile is a two-stage build. A Fedora 42 builder stage runs dnf --installroot into /target, applies transforms, runs chroot operations, and validates. A FROM scratch final stage copies the assembled rootfs.

COPYBUILDER STAGEfedora:42 AS builderdnf —installroot255 packages into /target7 transformssymlinks, /var, nss-altfiles, dracutchroot ops, ostree-boot, rpmdbbootc container lint11 of 12 checks passFINAL IMAGEFROM scratchCOPY —from=builder /target/ /LABEL containers.bootc 1CMD [“/sbin/init”]IN THE IMAGEOS corekernel, systemd, bootc, ostree, SELinuxIN THE IMAGERuntime and networkpodman, buildah, skopeo, networkd, nftables, sshdNOT IN THE IMAGENo build toolingno rpm-ostree, treefile, inheritanceRESULT886 MB255 packages, lint 11 of 12

Everything that builds the image stays in the builder stage. The scratch stage is three directives over a rootfs that was already assembled, transformed and linted, so the 886 MB that ships carries a kernel, an init, a container runtime and a network stack, and nothing that built it.

FROM quay.io/fedora/fedora:42 AS builder
ARG RELEASE=42

RUN dnf --use-host-config \
    --installroot=/target \
    --releasever=${RELEASE} \
    --setopt=install_weak_deps=False \
    --nodocs \
    -y install \
    kernel systemd systemd-pam bootc bootupd \
    ostree nss-altfiles selinux-policy-targeted \
    # ... all 255 packages

--use-host-config tells dnf to use the builder’s repo configuration for the target installroot. --setopt=install_weak_deps=False prevents recommended packages from inflating the install. --nodocs saves space.

After the packages are installed, the rootfs needs seven categories of transforms to become a valid bootc image.

SEVEN TRANSFORMS, 1 TO 4 FILESYSTEM LAYOUT, 5 TO 7 TOOLING AND BOOTostree layout, checkedby bootc container lintconf.d must exist beforedracut runs in step 5/root becomes a real dirfor dracut, symlink afterPR #2100 folds steps 1 to 4and 6 to 7 into bootccontainer finalize-rootfs1Toplevel symlinks/home to var/home, /root to var/roothome, /ostree to sysroot2/var cleanup and tmpfiles.dgenerate tmpfiles.d entries, clean transient /var state3nss-altfiles and /usr/lib/passwdcopy /etc/passwd to /usr/lib/passwd for an immutable /etc4Dracut configurationwrite conf.d before dracut runs, fix the /root symlink5Chroot operationsmount /proc /sys /dev, run dracut, preset-all, bootupd6/usr/lib/ostree-bootcopy the EFI files from /boot into /usr/lib/ostree-boot7rpmdb relocationmove the rpm database, inject the config files

Seven transforms turn a plain dnf installroot into a valid bootc image, and the order carries weight: dracut in step 5 reads the conf.d written in step 4. Step 5 is dashed because it mounts /proc, /sys and /dev and chroots in, the one operation CI cannot do unprivileged.

Seven things that broke

ostree expects /home to be a symlink to var/home, /root to var/roothome, /opt to var/opt, and so on. dnf --installroot creates real directories. The fix is straightforward:

rm -rf /target/home && ln -s var/home /target/home
rm -rf /target/root && ln -s var/roothome /target/root
rm -rf /target/opt && ln -s var/opt /target/opt
# ... and /srv, /mnt, /media, /usr/local
mkdir -p /target/sysroot
ln -s sysroot/ostree /target/ostree

2. /var cleanup and tmpfiles.d

ostree manages /var as mutable state created at boot via tmpfiles.d. Content in /var at build time that does not have a corresponding tmpfiles.d entry triggers a lint warning. The fix: generate tmpfiles.d entries for the directories you need and clean transient state from /var.

3. nss-altfiles and /usr/lib/passwd

nss-altfiles reads system accounts from /usr/lib/passwd and /usr/lib/group on immutable systems where /etc may be regenerated. The RPM installs the NSS module but not the static files. rpm-ostree compose generates these from reference files. With dnf --installroot, you copy them manually:

cp /target/etc/passwd /target/usr/lib/passwd
cp /target/etc/group /target/usr/lib/group

We validated at runtime that nss-altfiles correctly serves system accounts from /usr/lib/passwd while new users created with useradd go to /etc/passwd. getent passwd root resolves through the altfiles path, getent passwd testuser resolves through the local path. Both work.

4. Dracut configuration ordering

The ostree and bootc dracut modules are only included in the initramfs if configuration files requesting them exist before dracut runs. Write the conf.d files first, then run dracut:

# Must exist BEFORE dracut runs
cat > /target/usr/lib/dracut/dracut.conf.d/20-bootc-base.conf << 'EOF'
hostonly=no
add_dracutmodules+=" kernel-modules dracut-systemd systemd-initrd base ostree "
EOF

5. /proc, /sys, /dev for chroot operations

dracut, systemctl preset-all, and bootupctl backend generate-update-metadata all need /proc, /sys, and /dev mounted in the chroot. A plain chroot /target dracut ... without these mounts fails silently or with opaque errors.

6. /usr/lib/ostree-boot

bootupd expects /usr/lib/ostree-boot/ to contain EFI bootloader files for metadata generation. In the ostree build path, this directory is populated by the treefile postprocess. With dnf --installroot, the bootloader files land in /boot/efi/ and /boot/grub2/. Copy them to /usr/lib/ostree-boot/ before running bootupd, then clean /boot for the final image.

7. NBD for install-to-disk

bootc install to-disk needs a real block device with partition support. Loop devices (losetup) do not reliably expose partition devices inside containers. NBD (qemu-nbd) does:

modprobe nbd max_part=8
qemu-nbd --connect=/dev/nbd0 disk.qcow2
bootc install to-disk --filesystem xfs --wipe /dev/nbd0
qemu-nbd --disconnect /dev/nbd0

This is not a build issue per se, but it blocked validation until we switched from loop devices to NBD. Our validation harness already used this pattern; we just forgot to apply it here.

The hardening phase

The initial image booted and passed basic smoke tests. Then we asked: is this actually ready to be a production base? The answer was no. We identified 11 categories of validation that a real container host needs and tested each one.

What the initial image was missing

The first build had 236 packages and passed bootc container lint, but:

  • No ip command. The iproute package was not installed. You could not run ip addr or ss -tlnp on the system.
  • No pager. less was missing. journalctl output scrolled off the screen.
  • No tar. Could not extract archives.
  • No editor. vim-minimal was not installed. No vi on the system at all.
  • No timezone data. tzdata was missing. timedatectl set-timezone America/New_York returned “Invalid or not installed time zone.”
  • Volatile journal. No Storage=persistent in journald.conf. Logs were lost on every reboot. After three reboots, only the current and previous boot were visible in journalctl --list-boots.

These are the kind of gaps that make a system hostile to operate in an emergency. SSH in, can’t check the network, can’t read logs, can’t edit a file, can’t set the timezone. The system boots and runs containers, but the operator experience is broken.

What we learned from the upstream tiers

The upstream tier structure told us exactly what we were missing. The minimal-plus tier — which is the shared base for all Fedora image variants including IoT, Atomic Desktops, and CoreOS — includes iproute, less, tar, vim-minimal, tzdata, attr, hostname, and bash-completion. These are not luxury packages. They are the minimum for an administrable system.

The persistent journal configuration is only in the standard tier. This means anyone building from minimal or minimal-plus gets volatile journal by default. This is arguably a bug in the upstream tier design — a server that loses its logs on reboot is not production-ready regardless of how minimal it is.

UPSTREAM FEDORA BOOTC TIERSWHAT OUR 255 PKGS TAKEstandardfedora-bootc:42, 523 packagesadds NetworkManager, SSSD, sos, man-db,recommends: true= minimal-plus + thispersistent journalminimal-plusshared base for IoT, Atomic, CoreOSadds podman, openssh, iproute, less,vim-minimal, tar, tzdata, sudo, polkit= minimal + thisminimalthe floorkernel, systemd, bootc, dnf, SELinuxthe journal setting,nothing elseevery essentialon this rungthe whole rungBuild below standard and journald stays volatile: the logs go away at the next reboot.

Each upstream tier is the one below it plus a list. The operator essentials wait for minimal-plus and persistent journal for standard, so a build that stops short loses its logs at reboot. We take both and skip the rest of standard.

The 11 validation categories

After adding the missing packages and the journal configuration, we re-validated across every category:

CategoryWhat we testedResult
Package auditAll critical CLI tools present (ip, ss, less, tar, vi, jq, ausearch)PASS
Journal persistence5 reboots, all previous boots visible in journalctl --list-bootsPASS
User managementuseradd, nss-altfiles runtime, sudo, su, survives rebootPASS
bootc upgradebootc switch to live registry image, reboot, rollback, reboot backPASS
Rootless podmanService user with subuid/subgid, linger, pull, run, long-lived containerPASS
Console accessgetty on tty1, serial-getty on ttyS0, login prompt in serial logPASS
Locale/timezoneC.UTF-8 functional, timedatectl set-timezone works, tzdata presentPASS
SELinux deepTargeted policy, enforcing, 366 booleans, correct file contexts, zero AVCsPASS
Essential CLI toolsip, ss, ps, top, mount, lsblk, less, tar, vi, jq, curl all presentPASS
tmpfiles.dAll /var directories provisioned at boot including /var/log/journalPASS
ZFS module layerMulti-stage kernel module build layers on top without modificationPASS

The bootc upgrade test

The upgrade test was the most important validation. We ran bootc switch from our minimal-base to the published basef image, rebooted into it (768 packages, completely different base), then rolled back to our minimal-base and rebooted again. Both transitions were clean. The ostree deploy machinery does not care whether the image was built from fedora-bootc:42 or from scratch — it deploys the OCI layers the same way.

The ZFS module layer test

The multi-stage kernel module pattern compiles ZFS from source in a throwaway builder stage and copies the compiled modules into the final image via a manifest. We built the ZFS layer on top of our minimal-base using the exact same Containerfiles from the production pipeline. No modifications needed. depmod, modinfo, ldd, systemctl enable, and bootc container lint all passed.

The final image

After transforms, the FROM scratch stage is two lines:

FROM scratch
COPY --from=builder /target/ /
LABEL containers.bootc 1
CMD ["/sbin/init"]

bootc container lint passes 11 of 12 checks (1 skipped, cosmetic warnings only). Zero failures.

Size comparison

minimal-basethis image886 MB, 255 pkgsbase-zfson minimal-base1.89 GB, 368 pkgsfedora-bootc:42upstream standard2.1 GB, 523 pkgsbasef-baseon fedora-bootc:422.3 GB, 540 pkgsbase-zfson fedora-bootc:422.8 GB, 650 pkgs

The minimal base is 58 percent smaller than upstream, and the saving survives layering: the ZFS image built on it is a third smaller than the same layer on fedora-bootc:42.

The minimal base is 58% smaller than upstream. The ZFS layer on minimal-base is 32% smaller than the current ZFS layer on fedora-bootc:42. Every package in the image is a conscious choice.

What this means

The dnf --installroot path works for building bootc images from scratch. The seven transforms documented here are the complete set needed to go from a plain dnf rootfs to a bootable, lint-passing bootc image. When PR #2100 merges upstream, most of these transforms collapse into bootc container finalize-rootfs.

The practical implication: you do not need to inherit from fedora-bootc:42 or centos-bootc:stream10. You do not need rpm-ostree compose or a treefile. A standard Containerfile with dnf --installroot and ~20 lines of filesystem transforms produces a bootable image that passes all the same validation checks.

The hardening phase taught us that “boots and runs containers” is necessary but not sufficient. A production base image needs persistent journal, timezone data, network diagnostic tools, an editor, and a pager. The upstream minimal tier is missing several of these. The minimal-plus tier gets closer but still lacks persistent journal. Building from scratch means you own these decisions explicitly rather than inheriting them from a tier that may or may not match your needs.

Building this in CI: rootless Podman and buildah isolation

The Containerfile worked locally. When we merged it into the basef production pipeline, the CI build failed immediately.

Our CI runner is a rootless Podman setup: a GitLab Runner in a Podman quadlet container (AddCapability=all, SecurityLabelDisable=true), using the Docker executor talking to the rootless Podman socket. The runner config sets privileged = true and BUILDAH_ISOLATION=chroot. CI job containers get --privileged from rootless Podman.

The Containerfile uses mount -t proc proc /target/proc inside RUN commands for chroot operations (dracut, systemctl preset-all, bootupd). This requires CAP_SYS_ADMIN. The debugging path to figure out why it was denied took six attempts.

The debugging path

1. First failure: mount denied. mount -t proc proc /target/proc: permission denied. BUILDAH_ISOLATION=chroot denies mount(2) even in privileged containers. Buildah’s chroot isolation restricts RUN sub-processes regardless of the outer container’s privilege level.

2. Bind mounts. Tried mount --bind /proc /target/proc. Also denied — same restriction. Chroot isolation blocks all mount operations in RUN commands, not just new filesystem mounts.

3. Avoid chroot entirely. Tried dracut --sysroot /target and systemctl --root=/target to skip the chroot. dracut ran but failed: “Module ‘ostree’ cannot be installed.” The --sysroot mode uses the builder’s dracut modules, but the ostree dracut module only exists in the target rootfs. The builder (plain fedora:42) does not have it.

4. Install dracut on the builder. Installed dracut and systemd on the builder stage. dracut found the binary but still failed: /dev/kmsg: Permission denied and the ostree module still could not be found because --sysroot looks for modules on the host, not in the sysroot.

5. Root cause: tested on the runner host directly.

# Direct container — mount works
podman run --privileged fedora:42 \
  mount -t proc proc /tmp/p
# ✓ OK

# Buildah chroot isolation — mount denied
podman run --privileged fedora:42 \
  buildah build --isolation=chroot . \
  # RUN mount -t proc proc /tmp/p
# ✗ permission denied

# Buildah OCI isolation — also denied
podman run --privileged fedora:42 \
  buildah build --isolation=oci . \
  # RUN mount -t proc proc /tmp/p
# ✗ permission denied (nested OCI can't mount in rootless)

# Buildah with explicit cap-add — mount works
podman run --privileged fedora:42 \
  buildah build --cap-add=SYS_ADMIN . \
  # RUN mount -t proc proc /tmp/p
# ✓ OK

The key: --privileged gives capabilities to the container process, but buildah does not pass them to RUN sub-processes unless you explicitly use --cap-add=SYS_ADMIN on the buildah build command. The capabilities of the outer container and the capabilities available inside a buildah RUN step are independent.

6. Fix: one line per build job. Added --cap-add=SYS_ADMIN to the buildah build commands in the CI/CD pipeline component. Pipeline went fully green.

The capability chain

The full privilege path for a FROM-scratch bootc build in rootless Podman CI:

ROOTLESS HOSTpodmanuser namespaceRUNNERcontainer—privilegedCI JOBcontainer—privilegedRESTRICTED BY DEFAULTbuildah build RUN—cap-add=SYS_ADMIN unlocks mount—privileged does not propagate into buildah RUN sub-processes

Privilege stops at the job container. The buildah RUN step runs under its own capability set, so a FROM-scratch build that needs mount(2) has to be granted SYS_ADMIN on the buildah command itself.

This interaction is undocumented. BUILDAH_ISOLATION=chroot is the default for many CI setups and is the right default for standard Containerfiles. But FROM-scratch bootc images need mount(2) for chroot operations, which chroot isolation blocks regardless of the outer container’s privilege level.

The security tradeoff is real: --cap-add=SYS_ADMIN grants mount capability to arbitrary Containerfile RUN commands. For a self-hosted runner building your own images, this is acceptable. For shared runners, consider a dedicated build host with appropriate isolation.

For container hosts where every application runs in Podman, this matters. The base OS can be genuinely minimal: kernel, systemd, container runtime, SSH, firewall, and the admin tools you actually need. Nothing else. The multi-stage kernel module pattern layers cleanly on top. The base got smaller, but the layering pattern is unchanged.