RSDK-14136: Docker base image CI - #670
Conversation
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
| - name: Create multi-arch manifest | ||
| run: | | ||
| # imagetools create -t <tag> <image>@sha256:<d1> <image>@sha256:<d2> | ||
| docker buildx imagetools create -t "${{ matrix.cell.tag }}" \ |
There was a problem hiding this comment.
I was thinking about this step as "doing the merge", and then went looking for where we "publish" or "push" the merge, but I didn't see such a step. Is it implied here? If so, maybe worth a comment.
There was a problem hiding this comment.
maybe addressed by the newly added comment?
There was a problem hiding this comment.
Not entirely. The way the steps here read to me are:
- login
- download the artifacts matching a pattern
- run
imagetools createto merge.
But I don't see anything that "pushes" or "publishes" the final merged images to ghcr. I'm just curious where it happens, since it is non-obvious to me.
There was a problem hiding this comment.
ah ok yes good question, will add some comments to that effect. actual image pushing happens if --push=true is passed to docker buildx bake in the build job, and then docker buildx imagetools create pushes the manifest list in the merge job
|
|
||
| - name: Create multi-arch manifest | ||
| run: | | ||
| # imagetools create -t <tag> <image>@sha256:<d1> <image>@sha256:<d2> |
There was a problem hiding this comment.
This comment is maybe not very useful, and I don't know what d1 and d2 are, actually. Maybe it is worth explaining, but as a prose comnent.
| | { target: .key, tag: (.value.tags[0]), image: (.value.tags[0] | split(":")[0]) } | ||
| ]') | ||
| echo "cells=$cells" >> "$GITHUB_OUTPUT" | ||
| echo "$cells" | jq . |
There was a problem hiding this comment.
Is this second echo just for debugging?
There was a problem hiding this comment.
yep! added a comment, now in the shell script
| # Each system-* target -> { target, tag, image }. `tag` is the published | ||
| # multi-arch reference the merge job creates; `image` is that tag with | ||
| # the version stripped (the per-cell repo digests are pushed under). | ||
| cells=$(docker buildx bake -f docker-bake.hcl system --print \ |
There was a problem hiding this comment.
Arguably is reaching the level of complexity where it should be a shell script in the workflow directory.
| run: | | ||
| digest=$(jq -r '.["${{ matrix.cell.target }}"]."containerimage.digest"' metadata.json) | ||
| test -n "$digest" && test "$digest" != null | ||
| mkdir -p /tmp/digests |
There was a problem hiding this comment.
The use of digest files on disk in tmp feels like there should be a better way. We are uploading all the digests as artifacts, and the uploads follow a pattern. And we download them by pattern too it looks like? I think this use of tmp makes it look like some state is flowing through /tmp stage to stage but it maybe really isn't? Or maybe I'm misinterpreting it or misunderstanding what it is really doing. Anyway, it confused me a bit so let's look at it.
There was a problem hiding this comment.
yeah it is just being used as per-runner scratch space, cross-job state is built and stored as artifacts which are then merged below. refactored this a bit and added some comments on "magic" lines
| pattern: digest-${{ matrix.cell.target }}-* | ||
| merge-multiple: true | ||
|
|
||
| - name: Create multi-arch manifest |
There was a problem hiding this comment.
I went and looked at the artifacts as seen on the github packages page. Some notes:
- The
cpp-sdk-system-distro:versionsetup is great. Super clear when you pull up https://github.com/viamrobotics/viam-cpp-sdk/pkgs/container/cpp-sdk-system-debian and you see all the debians there. - The packages are multiarch! I was able to run
docker run --platform linux/arm64 --rm -it ghcr.io/viamrobotics/cpp-sdk-system-debian:trixieto get an arm64 run and then change toamd64. Awesome. - It appears that
bookwormhas somehow becomelatest: https://github.com/viamrobotics/viam-cpp-sdk/pkgs/container/cpp-sdk-system-debian/991778263?tag=bookworm? I think this is github declaring it to be "latest", not a docker tag, because if I run like-it ghcr.io/viamrobotics/cpp-sdk-system-debianwithout a:xxx, it saysghcr.io/viamrobotics/cpp-sdk-system-debian:latest: not found. So, maybe this isn't something to worry about. On the other hand, it we DID want a latest tag, to what should it adhere? I'd guesstrixiefor debian andresoluteforubuntu? But, I'd be fine just not doing anything. If github is just using "date of publish" for marking something as latest, I don't think there is much we can do about it. - As noted above, the multiarch setup is great. I went to look at the
archpage though and saw something surprising. Along with the two expected archs, there is anunknown/unknown: https://github.com/viamrobotics/viam-cpp-sdk/pkgs/container/cpp-sdk-system-debian/991778263?tag=bookworm. I don't quite know how to figure out what that is, but it is probably worth investigating or at least understanding. - We should think a bit about versioning. We are going to want users of these to pin, so that they don't get auto-updated unless they want to be. The C++ SDK, for instance, will want to pin for itself, so that you can make a commit that updates the docker images, and then once they are built, another commit to move to them. You could just use the sha256 tags, but then you lose the arch independence. The version can't be just a tag, since then it doesn't identify the distro version. What shape does the version take? Incrementing? Date stamp? Where do we tack it into the label? Is it
ubuntu:focal-20260705? I'm open to suggestions here.
There was a problem hiding this comment.
apparently the "latest" tag is just a web-ui thing for the most recently uploaded package. we probably don't want to be pushing a latest tag, although it is doable. feel like that often bites us on CI jobs haha.
versioning...great question and was also wondering about that. there may not be org-level prior art here? i think that our containers (eg rdk-devenv) are not versioned or tend to be so rarely updated that it's kind of not an issue? an admirable level of stability to have, although realistically there are going to be speedbumps for the first little while after this goes live.
There was a problem hiding this comment.
I'm fine with getting this up and running without versioning, but I do think it will become important. While the base images are likely to be pretty stable, I expect they will have changes here and there.The conan images on the other hand I can see iterating more rapidly, as we add new profiles, or learn more about how the system really behaves.
Given that using the hash to pin undermines the multiarch nature of things, I think some versioning scheme is going to be needed: we really do want consumers to pin, so that we can update the containers without breaking them. Then consumers can update when they want or need to, in one coherent PR.
So, I do think coming up with some sort of versioning plan is needed. My suggestion would be a new JIRA ticket in the epic, so we don't lose track of that goal.
I'm also still curious what is up with that unknown/unknown in the arch.
There was a problem hiding this comment.
ok that was pretty interesting, just checked with claude and unknown/unknown is an unrunnable sentinel value for os/arch used for the SLSA provenance attestation manifest/SBOM because buildkit attaches these manifests/attestations by default
There was a problem hiding this comment.
https://viam.atlassian.net/browse/RSDK-14221 for the versioning task, i have a claude session on it now
There was a problem hiding this comment.
Great, I will be interested to see what it comes up with, as I don't actually know the right answer either.
The SLSA provenance thing is fascinating, and also explains why I was unable to run the image by hash, which I indeed tried to do.
There was a problem hiding this comment.
It appears that it being displayed by GH is viewed at least by some as a bug: https://github.com/orgs/community/discussions/45969
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Andrew C. Morrow (acmorrow)
left a comment
There was a problem hiding this comment.
LGTM, though:
- I think a new versioning ticket in the epic is desirable
- I'm still a little confused about where the "push" / "publish" things are happening. Perhaps they are implciit? If so, some commentary that this is actually building TO ghcr is needed.
- I'd like a clarification on what the
unknown/unknownarchs are all about.
| - name: Create multi-arch manifest | ||
| run: | | ||
| # imagetools create -t <tag> <image>@sha256:<d1> <image>@sha256:<d2> | ||
| docker buildx imagetools create -t "${{ matrix.cell.tag }}" \ |
There was a problem hiding this comment.
Not entirely. The way the steps here read to me are:
- login
- download the artifacts matching a pattern
- run
imagetools createto merge.
But I don't see anything that "pushes" or "publishes" the final merged images to ghcr. I'm just curious where it happens, since it is non-obvious to me.
| pattern: digest-${{ matrix.cell.target }}-* | ||
| merge-multiple: true | ||
|
|
||
| - name: Create multi-arch manifest |
There was a problem hiding this comment.
I'm fine with getting this up and running without versioning, but I do think it will become important. While the base images are likely to be pretty stable, I expect they will have changes here and there.The conan images on the other hand I can see iterating more rapidly, as we add new profiles, or learn more about how the system really behaves.
Given that using the hash to pin undermines the multiarch nature of things, I think some versioning scheme is going to be needed: we really do want consumers to pin, so that we can update the containers without breaking them. Then consumers can update when they want or need to, in one coherent PR.
So, I do think coming up with some sort of versioning plan is needed. My suggestion would be a new JIRA ticket in the epic, so we don't lose track of that goal.
I'm also still curious what is up with that unknown/unknown in the arch.
CI job to build, merge, and publish the recently added docker images
See
https://github.com/viamrobotics/viam-cpp-sdk/pkgs/container/cpp-sdk-system-ubuntu
https://github.com/viamrobotics/viam-cpp-sdk/pkgs/container/cpp-sdk-system-debian
for successful non-dry runs