Eight years ago
Akihiro Suda's buildbench compared Docker, BuildKit, img, Buildah and kaniko in 2018. kaniko came out last on the rebuild 15x slower than buildkit. That number still shapes how people think of kaniko.
In 2024 we forked kaniko with the goal to modernize it and bring it on par with buildkit. Most of that effort was on Dockerfile syntax support, but recently we also looked into performance improvements. Here I would like to share what we achieved so far and where we're heading next.
The benchmark
We reuse buildbench's ex01 Dockerfile however we do change the approach slightly. The original comparison was unfair because --cache-copy-layers was not activated on kaniko, foregoing half of the caching features and buildkit was allowed to use a local disk cache. For a fair comparison we use both tools with their best configuration in a CI build, this means only remote cache is allowed.
FROM alpine AS buildc
RUN apk add --no-cache build-base
RUN echo -e "#include <stdio.h>\nint main(int ac, char *av[]){printf(\"hello c\\\\n\");return 0;}" | tee /hello.c
COPY . /foo
RUN gcc -o /a.out /hello.c
# the COPY above SHOULD NOT invalidate the cache for the buildgo stage.
FROM alpine AS buildgo
RUN apk add --no-cache build-base
RUN apk add --no-cache go
RUN echo -e "package main\nfunc main(){println(\"hello go\")}" | tee /hello.go
RUN go build -o /a.out /hello.go
FROM alpine
COPY --from=buildc /a.out /hello1
COPY --from=buildgo /a.out /hello2
# Note: as of June 5, 2018, Buildah and Kaniko does not support `FROM anotherstage`
We build it in three situations.
- Cold. Nothing is cached yet, every command runs, every layer is pushed. This measures raw execution and the cost of writing the image and the cache. Most of the time goes into downloading packages, which no builder can speed up, so this is the floor they all share.
- Unchanged. The exact same inputs are built again. In CI this happens far more often than people expect. The ideal result here is close to zero work. This measures how much a builder has to fetch and unpack just to find out that nothing changed.
- Edit. A file in the context changes, as in every development iteration.
buildchas to rebuild from itsCOPYonwards,buildgois untouched. This measures cache granularity: whether a builder rebuilds only what the change reaches, or drags the independent stage along.
Together the three cases separate the cost of doing work, the cost of proving there is no work, and the cost of doing only the right work.
Results
| cold | unchanged | edit | |
|---|---|---|---|
| BuildKit v0.33.0 | 37.4 ± 2.5 s | 6.0 ± 0.1 s | 13.2 ± 1.5 s |
| Buildah v1.43.4 | 62.3 ± 6.7 s | 28.9 ± 10.0 s | 23.7 ± 2.3 s |
| kaniko v1.24.0 (Google) | 55.0 ± 3.4 s | 39.8 ± 0.7 s | 41.4 ± 2.4 s |
| kaniko v1.25.19 (Chainguard) | 50.8 ± 3.4 s | 40.1 ± 1.9 s | 40.6 ± 1.9 s |
| kaniko v1.29.0 (osscontainertools) | 45.7 ± 1.6 s | 9.6 ± 1.5 s | 26.7 ± 1.3 s |
Chainguard's fork, v1.25.19, performs like Google's v1.24.0 in all three cases, so everything said about v1.24.0 here applies to it as well.
Mean ± standard deviation of five repetitions on gitlab.com's hosted runners, all builders on the same runner, in random order, pushing to registry.gitlab.com. Every build starts without local state, only the registry cache carries over. kaniko v1.29.0 is v1.28.5 with the Preview profile, whose flags become the default in v1.29.0. Chainguard and Buildah were measured in separate runs under the same conditions. Buildah runs as a privileged container with native overlay storage.
Why the unchanged build got fast
Google's kaniko could already predict whether it had to do any work. Before running a command it computed the command's cache key and looked it up in the registry, and when every command of a build was a hit it skipped the build entirely. But kaniko was designed for single-stage builds. Multi-stage builds were grafted onto it later, and that is also where most of the bugs were when we forked it.
The lookahead did not work across stage boundaries. The cache key of COPY --from=buildc /a.out /hello1 was computed from the files it copies, and those files only exist once buildc has been built. So to find out whether the final stage had any work to do, kaniko had to build every stage before it. In ex01 that means pulling and unpacking the C toolchain, the Go toolchain and both binaries from the cache, only to learn that nothing changed. That is where Google's 40 seconds on the unchanged build go.
With v1.29.0 kaniko can now predict cache hits a priori for all stages. This means if your build is a 100% hit, all work gets eliminated. The builder stages disappear, and the final image is assembled from layers that are already in the registry. The 10 seconds that remain are mostly cache lookups and the push.
# kaniko v1.24
FROM alpine AS buildc
RUN apk add --no-cache build-base # hit
RUN echo ... | tee /hello.c # hit
COPY . /foo # hit
RUN gcc -o /a.out /hello.c # hit
FROM alpine AS buildgo
RUN apk add --no-cache build-base # hit
RUN apk add --no-cache go # hit
RUN echo ... | tee /hello.go # hit
RUN go build -o /a.out /hello.go # hit
FROM alpine
COPY --from=buildc /a.out /hello1 # hit, once buildc ran
COPY --from=buildgo /a.out /hello2 # hit, once buildgo ran
# kaniko v1.29.0
# buildc not built
# buildgo not built
FROM alpine
COPY --from=buildc /a.out /hello1 # hit
COPY --from=buildgo /a.out /hello2 # hit
Where BuildKit still wins
BuildKit is still ahead in all three cases, and in each one we know why.
Edit: a cached stage built for nothing
After the edit buildc rebuilds from its COPY onwards, and buildgo is fully cached. The final stage copies from both, but the mean part is that buildgo gets copied after buildc. So our work to do optimizations a priori can't be realised. Whether we have a cache hit on buildgo is only known after we have built buildc, as the key of COPY --from=buildc depends on what the rebuilt buildc produces. BuildKit on the other hand schedules work lazily, so it can skip the buildgo stage entirely here.
# kaniko v1.29.0
FROM alpine AS buildc
RUN apk add --no-cache build-base # hit
RUN echo ... | tee /hello.c # hit
COPY . /foo # miss
RUN gcc -o /a.out /hello.c # miss
FROM alpine AS buildgo
RUN apk add --no-cache build-base # hit
RUN apk add --no-cache go # hit
RUN echo ... | tee /hello.go # hit
RUN go build -o /a.out /hello.go # hit
FROM alpine
COPY --from=buildc /a.out /hello1 # hit, once buildc ran
COPY --from=buildgo /a.out /hello2 # hit, once buildc ran
# BuildKit
FROM alpine AS buildc
RUN apk add --no-cache build-base # hit
RUN echo ... | tee /hello.c # hit
COPY . /foo # miss
RUN gcc -o /a.out /hello.c # miss
# buildgo not built
FROM alpine
COPY --from=buildc /a.out /hello1 # hit
COPY --from=buildgo /a.out /hello2 # hit
We plan to cover that gap with a mixed approach, where the decision to build a stage can be deferred, but work is still planned ahead. Switching to a lazy evaluation model would give up a lot of control over the order of execution and we see a lot of potential in the ability to schedule work smartly.
Cold: a layer built, then downloaded again
Both builder stages start with the same RUN apk add --no-cache build-base. kaniko runs it in buildc, pushes the layer to the cache and moves on to buildgo. There it finds that layer in the cache, so it downloads the layer it uploaded a few seconds earlier and unpacks it, which takes longer than running apk again.
BuildKit sees that the step is identical in both stages, runs it once and reuses the result locally. That is its build graph at work: in LLB every step is a vertex identified by a digest of its definition and its inputs, so the same step in two stages is the same vertex.
Unchanged: assembling an image that already exists
In the fully cached case kaniko will still build the image locally from cache, it should realize that the image it would assemble together already exists in the registry and no work needs to be done.