add scripts/pr_lgtm_scanner.py — LGTM/KO scanner for pr-iphone-deploy

Standalone Python scanner extracted from the inline python in wait-approval. Embedding multi-line Python in a YAML run: | block scalar is fragile: YAML dedents block-scalar content to the first line's indent, which produced both YAML parse failures and Python IndentationErrors in earlier iterations of this workflow. A standalone script file sidesteps all of that and is testable locally.

Behavior preserved:
  - json.loads + \bLGTM\b / \bKO\b word-boundary regex (case-insensitive)
  - skip bot comments tagged with <!-- tabatago:ready-to-test -->
  - KO checked BEFORE LGTM (PR with both signals blocks)

Verified locally across 8 scenarios (bot-only, LGTM, KO, both, empty, usage error, KOM/LGTMX non-match, case-insensitive).
This commit is contained in:
2026-07-20 09:34:28 +02:00
parent 47106ff281
commit 7196cce237

View File

@@ -0,0 +1,70 @@
#!/usr/bin/env python3
"""LGTM/KO scanner for the pr-iphone-deploy workflow's wait-approval job.
Reads a Gitea PR comments JSON array (as returned by
``/api/v1/repos/{owner}/{repo}/issues/{pr}/comments``) on argv[1] and prints
exactly one of: ``LGTM``, ``KO``, ``PENDING``.
Rules
-----
* The bot's own "Ready to test" comment body literally contains the words
``LGTM`` and ``KO`` as instructions to the reviewer — without filtering it
out, the workflow would self-merge ~30s after deploy. Bot comments are
tagged with the sentinel ``<!-- tabatago:ready-to-test -->`` and skipped.
* ``KO`` is checked BEFORE ``LGTM`` — a PR with both signals blocks (matches
the existing "KO blocks" intent).
* Word-boundary regex (``\\bLGTM\\b`` / ``\\bKO\\b``, case-insensitive) so
"Tested, LGTM!" counts and "KOM" / "LGTMX" do not.
This script lives in ``scripts/`` (not inline in the workflow) because
embedding multi-line Python inside a YAML ``run: |`` block scalar is fragile:
YAML dedents block-scalar content to the first line, which produces either a
YAML parse failure or a Python ``IndentationError`` depending on how the
body is indented. A standalone file sidesteps all of that.
"""
import json
import re
import sys
MARKER = "<!-- tabatago:ready-to-test -->"
LGTM_RE = re.compile(r"\bLGTM\b", re.IGNORECASE)
KO_RE = re.compile(r"\bKO\b", re.IGNORECASE)
def decide(comments_json: str) -> str:
"""Return 'KO', 'LGTM', or 'PENDING' for the given comments JSON payload."""
try:
comments = json.loads(comments_json or "[]")
except Exception:
comments = []
hit_lgtm = False
hit_ko = False
for comment in comments:
body = comment.get("body") or ""
if MARKER in body:
# Skip the bot's own "Ready to test" comment — its body mentions
# LGTM/KO as instructions and would otherwise self-trigger.
continue
if KO_RE.search(body):
hit_ko = True
if LGTM_RE.search(body):
hit_lgtm = True
if hit_ko:
return "KO"
if hit_lgtm:
return "LGTM"
return "PENDING"
def main() -> int:
if len(sys.argv) != 2:
print("usage: pr_lgtm_scanner.py <comments_json>", file=sys.stderr)
return 2
print(decide(sys.argv[1]))
return 0
if __name__ == "__main__":
sys.exit(main())