Eight years ago
Akihiro Suda's buildbench compared Docker, BuildKit, img, Buildah and kaniko in 2018. kaniko came out last: rebuilding an unchanged image took it 15x longer than BuildKit. That number still shapes how people think of kaniko.
In 2024 we forked kaniko with the goal of modernising it and bringing it on par with BuildKit. Most of that effort went into Dockerfile syntax support, but recently we have also looked into performance. Here I would like to share what we have achieved so far and where we are heading next.
The benchmark
We reuse buildbench's ex01 Dockerfile however we do change the approach slightly. The original comparison was unfair because --cache-copy-layers was not activated on kaniko, foregoing half of the caching features and buildkit was allowed to use a local disk cache. For a fair comparison we use both tools with their best configuration in a CI build, this means only remote cache is allowed.
FROM alpine AS buildc
RUN apk add --no-cache build-base
RUN echo -e "#include <stdio.h>\nint main(int ac, char *av[]){printf(\"hello c\\\\n\");return 0;}" | tee /hello.c
COPY . /foo
RUN gcc -o /a.out /hello.c
# the COPY above SHOULD NOT invalidate the cache for the buildgo stage.
FROM alpine AS buildgo
RUN apk add --no-cache build-base
RUN apk add --no-cache go
RUN echo -e "package main\nfunc main(){println(\"hello go\")}" | tee /hello.go
RUN go build -o /a.out /hello.go
FROM alpine
COPY --from=buildc /a.out /hello1
COPY --from=buildgo /a.out /hello2
# Note: as of June 5, 2018, Buildah and Kaniko does not support `FROM anotherstage`
We build it in three situations.
- Cold. Nothing is cached yet, every command runs, every layer is pushed. This measures raw execution and the cost of writing the image and the cache. Most of the time goes into downloading packages, which no builder can speed up, so this is the floor they all share.
- Unchanged. The exact same inputs are built again. In CI this happens far more often than people expect. The ideal result here is close to zero work. This measures how much a builder has to fetch and unpack just to find out that nothing changed.
- Edit. A file in the context changes, as in every development iteration.
buildchas to rebuild from itsCOPYonwards,buildgois untouched. This measures cache granularity: whether a builder rebuilds only what the change reaches, or drags the independent stage along.
Together the three cases separate the cost of doing work, the cost of proving there is no work, and the cost of doing only the right work.
Results
| cold | unchanged | edit | |
|---|---|---|---|
| BuildKit v0.33.0 | 37.4 ± 2.5 s | 6.0 ± 0.1 s | 13.2 ± 1.5 s |
| Buildah v1.43.4 | 62.3 ± 6.7 s | 28.9 ± 10.0 s | 23.7 ± 2.3 s |
| kaniko v1.24.0 (Google) | 55.0 ± 3.4 s | 39.8 ± 0.7 s | 41.4 ± 2.4 s |
| kaniko v1.25.19 (Chainguard) | 50.8 ± 3.4 s | 40.1 ± 1.9 s | 40.6 ± 1.9 s |
| kaniko v1.29.0 (osscontainertools) | 45.7 ± 1.6 s | 9.6 ± 1.5 s | 26.7 ± 1.3 s |
Chainguard's fork, v1.25.19, performs like Google's v1.24.0 in all three cases, so everything said about v1.24.0 here applies to it as well.
Mean ± standard deviation of five repetitions on gitlab.com's hosted runners, all builders on the same runner, in random order, pushing to registry.gitlab.com. Every build starts without local state, only the registry cache carries over. kaniko v1.29.0 is v1.28.5 with the Preview profile, whose flags become the default in v1.29.0. Chainguard and Buildah were measured in separate runs under the same conditions. Buildah runs as a privileged container with native overlay storage.
Cache lookahead
Google's kaniko could already predict whether it had to do any work. Before running a command it computed the command's cache key and looked it up in the registry, and when every command of a build was a hit it skipped the build entirely. But kaniko was designed for single-stage builds. Multi-stage builds were grafted onto it later, and that is also where most of the bugs were when we forked it.
Similarly, the lookahead did not work across stage boundaries. The cache key of COPY --from=buildc /a.out /hello1 was computed from the files it copies, and those files only exist once buildc has been built. So to find out whether the final stage had any work to do, kaniko had to build every stage before it. In ex01 that means pulling and unpacking the C toolchain, the Go toolchain and both binaries from the cache, only to learn that nothing changed. That is where Google's 40 seconds on the unchanged build go.
With v1.29.0 kaniko can now predict cache hits a priori for all stages. This means if your build is a 100% hit, all work gets eliminated. The builder stages disappear, and the final image is assembled from layers that are already in the registry. The 10 seconds that remain are mostly cache lookups and the push.
# kaniko v1.24
FROM alpine AS buildc
RUN apk add --no-cache build-base # hit
RUN echo ... | tee /hello.c # hit
COPY . /foo # hit
RUN gcc -o /a.out /hello.c # hit
FROM alpine AS buildgo
RUN apk add --no-cache build-base # hit
RUN apk add --no-cache go # hit
RUN echo ... | tee /hello.go # hit
RUN go build -o /a.out /hello.go # hit
FROM alpine
COPY --from=buildc /a.out /hello1 # hit, once buildc ran
COPY --from=buildgo /a.out /hello2 # hit, once buildgo ran
# kaniko v1.29.0
# buildc not built
# buildgo not built
FROM alpine
COPY --from=buildc /a.out /hello1 # hit
COPY --from=buildgo /a.out /hello2 # hit
Our next steps
BuildKit is still ahead in all three cases. But we know why, which makes them tractable.
Edit
After the edit, buildc rebuilds from its COPY onwards and buildgo is fully cached. The annoying part is the order: the final stage copies from buildc first and from buildgo second.
The key of COPY --from=buildc depends on what the rebuilt buildc produces, so it cannot be computed in advance. Our prediction stops at that command, and everything behind it in the stage falls back to being decided at build time, including the copy from buildgo. So buildgo gets built, even though nothing in it changed and the copy that follows would have been a hit either way. BuildKit schedules lazily and only materialises a stage once something that actually missed needs its output. Here nothing does.
# kaniko v1.29.0
FROM alpine AS buildc
RUN apk add --no-cache build-base # hit
RUN echo ... | tee /hello.c # hit
COPY . /foo # miss
RUN gcc -o /a.out /hello.c # miss
FROM alpine AS buildgo
RUN apk add --no-cache build-base # hit
RUN apk add --no-cache go # hit
RUN echo ... | tee /hello.go # hit
RUN go build -o /a.out /hello.go # hit
FROM alpine
COPY --from=buildc /a.out /hello1 # hit, once buildc ran
COPY --from=buildgo /a.out /hello2 # hit, once buildc ran
# BuildKit
FROM alpine AS buildc
RUN apk add --no-cache build-base # hit
RUN echo ... | tee /hello.c # hit
COPY . /foo # miss
RUN gcc -o /a.out /hello.c # miss
# buildgo not built
FROM alpine
COPY --from=buildc /a.out /hello1 # hit
COPY --from=buildgo /a.out /hello2 # hit
We are closing this with a mixed approach, where the decision to build a stage can be deferred but the work is still planned ahead. Going fully lazy would mean giving up control over the order of execution, and we see too much potential in scheduling work deliberately to trade that away.
Cold
Both builder stages start with the same RUN apk add --no-cache build-base. kaniko runs it in buildc, pushes the layer to the cache and moves on to buildgo. There it finds that layer in the cache, so it downloads the layer it uploaded a few seconds earlier and unpacks it, which takes longer than running apk again.
BuildKit sees that the step is identical in both stages, runs it once and reuses the result locally. That is its build graph at work: in LLB every step is a vertex identified by a digest of its definition and its inputs, so the same step in two stages is the same vertex.
Unchanged
Even when every command is a hit, kaniko still assembles the image: it pulls the cached layers, stacks them and pushes the result. But the image it is about to produce is already in the registry. That is precisely what the cache hits told it. Recognising that and stopping there would turn most of the remaining ten seconds into a handful of manifest lookups.