Skip to content

bug(gateway): provider environment revision is not deterministic when a provider profile has several annotations #3929

Description

@rain-sicoreai

User Story

As a platform operator driving the OpenShell gateway API (operator workspace mode, Kubernetes driver) to create sandboxes with bound provider credentials,
I want the provider environment revision to be stable while nothing changes,
so that the supervisor installs the provider credentials and policy updates apply promptly.

Problem Statement

When a provider profile has more than one entry in a map field (for example two annotations), the gateway reports a different provider environment revision on successive calls although no provider, profile or policy changed. In our runs GetSandboxProviderEnvironment alternated between exactly two provider_env_revision values from poll to poll for the same sandbox, and the configuration snapshot's revision did not match the environment's.

The revision hash includes the profile's protobuf encoding (hash_scoped_profile_revision in crates/openshell-server/src/provider_profile_sources.rs calls entry.response.encode_to_vec(); also lines 105 and 126 on main). ProviderProfile.annotations is map<string, string> and openshell-core's build does not configure ordered maps, so the encoded bytes depend on hash-map iteration order.

Impact / Why This Matters

  • On every settings poll the supervisor logs Settings poll: config change detected [... provider_env_changed:true], then CONFIG:FAIL_CLOSED [HIGH] Provider environment refresh failed; static credentials were revoked ... and Provider environment is unavailable or changed during preparation: provider credentials are withheld from the workload for as long as the sandbox runs.
  • ReportEndpointStatus fails with FailedPrecondition ("tool server endpoint status revisions do not match the current sandbox configuration") on every report.
  • Sandbox startup can fail with Startup configuration did not stabilize after 5 attempts, and a policy update took ~60 s to load instead of ~10 s, because preparation keeps being retried.
  • Workaround: keep at most one entry in every map field of provider profiles. With one annotation the revision is stable, credentials are installed and policy updates load within ~10 s.

Acceptance Criteria

  • The provider environment revision (config snapshot and GetSandboxProviderEnvironment) is identical across calls when providers, profiles and policy are unchanged, for profiles with several annotations (and any other map field).
  • A sandbox whose provider profile has several annotations installs its static credentials, with no FAIL_CLOSED provider refresh events and no ReportEndpointStatus FailedPrecondition in steady state.

Reproduction Steps

  1. Import a provider profile with two annotations, e.g. ImportProviderProfiles with annotations: {"example.com/a": "1", "example.com/b": "2"} and one credential.
  2. Create a provider of that type and a sandbox whose policy binds an endpoint to it (credential_binding).
  3. Watch the gateway log line GetSandboxProviderEnvironment request completed successfully ... provider_env_revision=<n> over a few minutes: the value alternates between two numbers.
  4. Watch the supervisor log: repeated provider_env_changed:true and CONFIG:FAIL_CLOSED events.
  5. Repeat with a single annotation: the revision is stable and the events stop.

Environment

  • OpenShell: gateway and supervisor v0.1.0 (Helm chart); the hashing code is unchanged on main as of 2026-09-30
  • OS: Ubuntu 24.04 (kernel 6.8)
  • Runtime: Kubernetes v1.34 (kubeadm), containerd 2.3; sandboxes on runc and on Kata Containers 4.2 (Cloud Hypervisor), same result
  • Deployment or integration: gateway in workspace_mode = "operator", driven through the gateway gRPC API by a control plane that creates sandboxes, provider profiles and providers

Logs

OCSF CONFIG:DETECTED [INFO] Settings poll: config change detected [old_revision:... new_revision:... policy_changed:false provider_env_changed:true]
OCSF CONFIG:FAIL_CLOSED [HIGH] Provider environment refresh failed; static credentials were revoked and previous dynamic grants remain active
OCSF CONFIG:CONFIGURATION_ERROR [HIGH] Provider environment is unavailable or changed during preparation
WARN openshell_supervisor::endpoint_status: Endpoint status report failed transiently; retaining immutable snapshot

Gateway, same sandbox, consecutive polls (no changes in between):

GetSandboxProviderEnvironment request completed successfully ... provider_env_revision=12427382202249825736
GetSandboxProviderEnvironment request completed successfully ... provider_env_revision=9706283966309050004
GetSandboxProviderEnvironment request completed successfully ... provider_env_revision=12427382202249825736

Activity

  1. letv1nnn commented on Sep 30, 2026

    @letv1nnn
    Contributor

    📋 triage-agent

    Triage Assessment

    Classification: validated-bug

    Summary

    Confirmed on main and in v0.1.0, the reporter's version, with high confidence. openshell-core generates protobuf map fields as std::collections::HashMap because crates/openshell-core/build.rs configures no btree_map. prost encodes map entries in iteration order. Every request re-decodes provider profiles from the store, so each decode builds a new HashMap with a fresh RandomState. The profile's encode_to_vec() output that feeds the provider environment revision therefore varies between calls whenever a map field has two or more entries.

    Investigation

    • Empirical reproduction on main. A ProviderProfile with two annotations, encoded once, then decoded and re-encoded 200 times, produced exactly 2 distinct byte encodings. This matches the reported alternation between two provider_env_revision values. With n entries, up to n! orderings are possible.
    • Hash sites that include profile encodings:
      • crates/openshell-server/src/provider_profile_sources.rs:105 and :126 (user profile source revision)
      • provider_profile_sources.rs:498, hash_scoped_profile_revision. Reached from grpc/policy.rs through compute_provider_env_revision_with_catalog_and_policy_bindings and compute_provider_env_revision_from_records_and_policy_bindings.
      • provider_profile_sources.rs:712 is test-only.
    • Already deterministic, not affected:
      • The provider's own credentials and credential_expiration_times, which policy.rs key-sorts before hashing.
      • The policy part of the revision, which uses deterministic_policy_hash in openshell-core/src/policy_identity.rs.
      • Only the profile encoding path lacks canonicalization.
    • Consistency within a call. Order is stable inside a single decoded instance, because HashMap::clone preserves hasher keys and layout. Each RPC builds its own catalog through snapshot_catalog, though, so GetSandboxConfig and GetSandboxProviderEnvironment independently land on one of the orderings and disagree about half the time. Separate gateway replicas would also disagree.
    • Supervisor reaction, which explains the logs:
      • openshell-supervisor/src/lib.rs sets provider_env_changed, requires the fetched environment identity to equal the desired identity, and otherwise emits the FAIL_CLOSED "static credentials were revoked" event and revokes the static environment.
      • Startup gives up after 5 unstable attempts.
      • On the server, configuration_generation_matches in grpc/policy.rs produces the ReportEndpointStatus FailedPrecondition.
    • Map fields affected beyond annotations. ProviderProfile.endpoints (sandbox.v1.NetworkEndpoint) carries nested maps: graphql_persisted_queries, L7Allow.query/params and L7DenyRule.query/params. Whether profile validation lets endpoints carry them was not checked; it affects scope, not validity. Profiles vended by interceptor sources go through the same :498 path.
    • Existing tests. No test uses a profile with more than one map entry, so the gap is untested.
    • Duplicates. None found. bug(supervisor): accept distinct provider environment revisions in process sidecar #2847 (closed) also involved provider_env_revision, but had a different cause: the process watcher treated the digest as a monotonic counter.

    Impact Signals

    • Affected users/scope: Any sandbox bound to a provider whose profile has two or more entries in any map field (annotations today, possibly nested endpoint maps). Static provider credentials are revoked on roughly half of settings polls, endpoint status reports fail, and startup can fail to stabilize. It fails closed: credentials are withheld, not exposed.
    • Regression: Unknown. The hashing is present in v0.1.0 and on main; no earlier known-good version was identified.
    • Workaround: Available but limiting: keep at most one entry in every map field of provider profiles.
    • Evidence quality: High. The code path was traced and the nondeterministic encoding was reproduced with a unit-level test.

    Candidate fix directions for whoever plans the work, not a plan: configure btree_map for the affected map paths in openshell-core/build.rs, which changes generated Rust types across crates and the Rust SDK, or canonicalize profile maps before hashing, reusing the sorted-map approach in policy_identity.rs. With either, a regression test with several annotations should assert that revisions are equal across decodes. Revision values will change once on upgrade, which triggers one benign provider-env refresh.

    Human Decision Required

    Decide whether OpenShell should address this issue. If yes, apply
    state:accepted, associate it with a roadmap item, or do both, and decide
    whether the work remains human-owned. Either action records acceptance;
    roadmap placement additionally records sequencing.
    To queue investigation or planning for an unattended agent, also apply
    agent:plan-requested. You can instead directly ask an agent to use
    create-spike or build-from-issue on this issue; the agent will warn about
    missing expected workflow labels and continue without changing them. If no,
    close it as not planned and record the rationale.

    Suggested labels (not applied, because the triaging account lacks triage permission): area:gateway, area:providers; replace state:triage-needed with state:validated.

  2. letv1nnn commented on Sep 30, 2026

    @letv1nnn
    Contributor

    hey @rain-sicoreai, I'd like to take this one

  3. ericcurtin commented on Oct 1, 2026

    @ericcurtin
    Contributor

    PR up: #4077. Happy to defer if you are already on it, @letv1nnn.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions