Repository navigation
Decode Python string escapes in PEP 508 marker literals (#19401) - #58
Merged
Merged
Conversation
Marker quoted strings now decode Python escapes (\n, \x41, A, \U0001F600, octal, etc.) in one pass, matching upstream's process_python_str/ast.literal_eval. Previously Token.Unquoted() returned the raw backslash sequence, so a requires_dist marker comparing against an escaped value compared against the wrong string. Validation is also tightened: truncated \u, truncated \U, and non-octal digits after a backslash (\8, \9) are now rejected, not silently accepted. Decoding happens once, in decodeQuotedStringContents (marker.go), at Unquoted()'s only production call site, so there is one source of truth for a literal's value; Unquoted() itself stays a raw, tokenizer-level accessor. This module's marker package is imported by PPM's src/pyresolve/target.go (marker.EnvironmentFromTarget) and evaluated during dependency resolution (PPM #20799, `get pypi --file-in --resolve-dependencies`), so this fixes a real marker-evaluation defect once a PPM release picks up a module version past the current v0.10.0 pin (PPM currently pins v0.9.0 via MVS, so this does not reach PPM immediately). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Decodes Python string-escape sequences in PEP 508 marker string literals, and tightens validation to reject truncated
\u/\Uescapes and invalid octal digits that were previously accepted.Stale issue claim — corrected
#19401 says "Product impact today: none. PPM consumes only this module's
license/package." That is false. PPM'ssrc/imports this module'smarkerpackage atsrc/pyresolve/target.goviamarker.EnvironmentFromTarget, and marker evaluation ships on a customer path in PPM #20799 (get pypi --file-in --resolve-dependencies). The live defect: arequires_distmarker comparing against a string with an escape (e.g."\n", a Windows path with\\) compared against the raw backslash sequence, not the decoded value.This does not reach PPM immediately — PPM currently pins
go-python-packaging v0.9.0and MVS resolves there until a release is cut and PPM's pin moves.What changed
Upstream (
pypa/packaging) tokenizes quoted strings permissively (no backslash handling in the tokenizer) and decodes viaprocess_python_str(_parser.py, effectivelyast.literal_eval), convertingSyntaxError/ValueErrorintoInvalidRequirement: Invalid quoted string(_parser.py:379). This repo's tokenizer already matched the permissive rule (#18640); this PR adds the missing decode step and closes validation gaps (truncated\u, truncated\U, invalid octal) upstream also rejects.Implemented as a single left-to-right pass (
decodeQuotedStringContents,marker.go) — not a validate-then-decode split. A two-pass design is the exact bug class #18640's review caught late: an interior\\pair immediately followed byxgets misread as a truncated\xescape by a second, independent scan. The new decoder consumes exactly one escape unit per step and never re-enters the escape switch on a byte already consumed as part of a prior escape. Test caseC:\\x41(marker_test.go) locks this in: decodes to the 6 literal charactersC:\x41, notC:A.Handles:
\\ \' \" \n \t \r \0 \a \b \f \v,\xHH(2 hex digits),\uHHHH(4 hex digits),\UHHHHHHHH(8 hex digits, rejects > U+10FFFF), and Python octal\OOO(1-3 digits). Fully custom decoder, notstrconv.Unquote— Go's octal/\xsemantics and quote-escaping rules diverge from Python's in ways that made bendingstrconv.Unquotemessier than a ~60-line purpose-built decoder.Decode-site choice: decoding lives in
marker.go, not inToken.Unquoted(). Grepped the whole module for.Unquoted(callers: exactly one production call site (marker.go:322, insideparseMarkerVar); the rest are test call sites intokenizer_test.go.marker,requirement, andreqtxtreach markers only throughpep508.ParseMarker/ParseFullMarker/ParseRequirement, which funnel through that one call site — none callUnquoted()directly. With a single caller, there's no benefit to pushing Python string-literal semantics into the tokenizer layer, which otherwise knows nothing about Python string grammar.Token.Unquoted()'s behavior and doc comment are otherwise unchanged (still returns raw, quote-stripped text);TestUnquoted_DoesNotDecodeEscapeslocks this in.Two independent fixes, each verified RED before GREEN
Fix A — decoding (RED captured against a variant that validated but returned the raw undecoded string):
GREEN after the decoder:
TestDecodeQuotedStringContents_Accepts(21/21) andTestParseMarker_QuotedStringLiteralDecodesEscapespass.Fix B — tightened validation (RED captured against a variant with full decoding but relaxed
\u/\Uwidth checks and permissive\8/\9):GREEN after the tightened checks:
TestDecodeQuotedStringContents_Rejects(11/11) andTestParseMarker_QuotedStringRejectsMalformedUnicodeEscapepass.Neither temporary variant is in the committed diff.
Existing-caller safety
Checked pre-existing marker strings with backslashes elsewhere in the module (
grammar_conformance_test.go, from #18640) — all still decode/validate the same way after this change. Full module test suite is green, not justinternal/pep508.Verification
go test ./... -count=1 -timeout 900s— all packagesok, 0 failures.go test ./... -race— all packagesok, 0 failures.golangci-lint run ./...at CI's pinned v2.11.2 (confirmed against.github/workflows/ci.yml) — 0 issues.gofmt -lon the 4 changed files — empty.Scope
Touches only
internal/pep508/marker.go,internal/pep508/marker_test.go,internal/pep508/tokenizer.go(doc comments only — no rule changes),internal/pep508/tokenizer_test.go. Does not touch theURL/WStoken rules orrequirement.go(L3/#19402's territory).NEWS
No entry — library-internal, pre-GA, this repo has no NEWS/CHANGELOG (uses
chore(release):commits). Deliberately omitted, not forgotten.Addresses rstudio/package-manager#19401.
🤖 Generated with Claude Code