Files
trufflehog/scripts/test/detect_changed_detectors.sh
3a022f9d59 Automate corpora testing in CI (#4927)
* add detector corpora test workflow and script

* only run once per PR, make comment descriptive, add handling for manual runs to get PR issue number

* comment out types to see result on all commits

* uncomment types

* remove table from comment

* comment out types

* Phase 0: add explicit pipefail and capture trufflehog stderr

* Phase 1: differential diffing PR vs main

* DEMO: loosen Stripe regex (will revert)

* DEMO: loosen JDBC regex (will revert)

* Phase 1 fix: add --allow-verification-overlap, fix no-diff detection

The bench uses --no-verification, so the engine's overlap-path dedup
(which exists to protect verifiers from duplicate calls) adds noise
without value here — it causes shifts in unrelated detectors when only
one detector's regex changes. Pair --allow-verification-overlap with
--no-verification so each detector's regex behavior is measured
independently.

Also fix the false 'no diff vs main' claim that triggered when
NEW/REMOVED were zero but total counts differed.

* revert jdbc detector change

* Phase 2: detector scoping, new-detector handling, blast radius, status emoji

* DEMO: loosen JDBC + add fictional acmevault detector

* Phase 2 fix: harden corpus byte counting against early trufflehog exit

awk's END block doesn't run when trufflehog exits before draining stdin
(SIGPIPE kills awk first), leaving the bytes file empty and breaking the
step with a `$((TOTAL_BYTES + ))` syntax error. Read the file with a
default of 0 and validate it's an integer before arithmetic. Also fold
unzstd/jq stderr into STDERR_FILE so benign Broken pipe notices stay
out of CI logs.

Co-Authored-By: Claude Opus 4.7 <[email protected]>

* Phase 3a (1/3): add hack/extract-keywords for detector keyword introspection

Static AST parse of a detector package to extract the strings returned by
its Keywords() method. Used by the upcoming keyword-corpus builder to fan
out per-detector GitHub Code Search queries during the corpora bench.

AST-first because each detector lives in its own package; importing them
dynamically would require codegen or `plugin`. Falls back to a regex over
the function body, then a directory-wide grep, when AST resolution can't
statically resolve the return value (helper calls, build-tagged variants).

Co-Authored-By: Claude Opus 4.7 <[email protected]>

* Phase 3a (2/3): add Layer 1 keyword corpus builder + workflow integration

build_keyword_corpus.py queries GitHub Code Search for each changed
detector's pre-filter keywords and emits a zstd-compressed JSONL whose
shape matches the existing S3 corpus exactly: each line is
`{"provenance": {...}, "content": "<raw file content>"}`. The corpora
script's existing `unzstd | jq -r .content` pipe handles it unchanged —
provenance is descriptive only and never reaches trufflehog.

Rate-limit policy is header-driven: the search bucket's
X-RateLimit-Remaining and X-RateLimit-Reset headers gate every call,
with a 2.1s floor between requests as belt-and-suspenders. 403/429s
honor Retry-After or fall back to the reset window. Cap is 100 unique
results per detector, deduped on (repo, path, sha), with a per-keyword
sub-cap so one popular keyword can't starve the others.

A sidecar JSON reports per-detector fetch counts and a thin_l1 list of
detectors whose total returned results were zero (or whose keyword
extraction failed). The diff script reads it via a new
--keyword-corpus-meta arg and renders a single contiguous blockquote
callout above the per-detector details — a sidecar instead of an
in-corpus signal because stdin metadata is dropped from trufflehog's
findings output.

Workflow change: a new "Build keyword corpus (Layer 1)" step fires after
detector detection and overwrites DATASETS via $GITHUB_ENV to append the
keyword corpus path. The corpora script picks it up unchanged through
its existing local-file branch.

Co-Authored-By: Claude Opus 4.7 <[email protected]>

* Phase 4 complete - Heatmap visualization

* Phase 4 rework (1/2): emit heatmap-grid.json sidecar from render_heatmap.py

Add a JSON sidecar that captures the same Δ matrix the PNG renders. The
diff script consumes this to render an emoji-bucketed Markdown table —
GitHub's PR-comment Markdown sanitizer strips data: URLs and serves
artifact zips behind auth, so neither inline base64 nor an artifact
<img src> embed actually displays. The PNG stays for artifact archival
and click-through.

Sidecar shape: {detectors, decoders, deltas, _layout}, with _layout
documenting the deltas[i][j] orientation inline so future readers don't
have to reverse-engineer it. Emitted whenever the grid is non-empty,
even if matplotlib import fails — the comment never depends on the PNG.

Co-Authored-By: Claude Opus 4.7 <[email protected]>

* Phase 4 rework (2/2): replace data-URL embed with emoji-bucketed Markdown table

GitHub's PR-comment Markdown sanitizer strips data: URLs from user
content, so the inline base64 PNG embed shipped in the prior commit
rendered as a broken link in the comment DOM (no <img> tag emitted).
Artifact-zip URLs require auth to download, so a fallback ![](...) link
is also a non-starter — it would render as a broken image.

Switch to a per-(detector, decoder) Δ table built from the grid JSON
sidecar render_heatmap.py now emits. Cells use emoji buckets aligned
with the existing status-emoji thresholds so the visual weight matches
the summary table:

  🟥 Δ ≥ +6   (matches NEW > 5 → 🔴)
  🟧 +1..+5
  ⬜ 0
  🟦 ≤ −1

Renders identically on web, mobile, email notifications, and CI log
replay — every surface a PR comment lands on. The colored matplotlib
PNG stays as a workflow artifact; when --heatmap-artifact-url is
supplied, the table is followed by a click-through link for reviewers
who want the rich version.

Workflow YAML mirrors the rename (--heatmap-png → --heatmap-grid) and
drops the base64-blob log filter — the report is back to plain
human-readable text.

Co-Authored-By: Claude Opus 4.7 <[email protected]>

* Phase 5 complete - Polish

* cleanup, enable verification

* fix bug

* optimizations

* cache keywords corpus

* rewrite comment message

* cache github api corpus per keyword

* cleanup

* remove github corpus

* revert changes for testing

* move Configure AWS credentials step to run only when detector changes are detected

* revert unnecessary changes

* cleanup + bugbot fixes

* run test with bigger (30gb) dataset, loosen jdbc regex

* optimizations

* bugbot fixes

* revert jdbc changes and bugbot fix

* run only on regex and/or keywords change

* bugbot fixes

* bugbot fix

* incorporate brad's comments, loosen jdbc regex to run a test to ensure everything works as expected

* revert test changes

* fix misleading bench skipped message

* run only once when PR opens

* pipe stderr directly to CI log instead of writing to file, loosen jdbc regex to trigger workflow

* testing: run on all commits

* add archive timeout

* remove bigger dataset

* revert testing changes

---------

Co-authored-by: Shahzad Haider <[email protected]>
Co-authored-by: Claude Opus 4.7 <[email protected]>
2026-05-15 14:47:53 +05:00

207 lines
7.4 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# detect_changed_detectors.sh — Phase 2
#
# Emits the list of detectors changed between two git refs, formatted for
# trufflehog's --include-detectors flag (comma-separated, lowercase protobuf
# enum names, optional ".v<n>" version suffix).
#
# Source of truth for each detector's identifier:
# - Proto enum name comes from the detector's Type() implementation in its
# source files (e.g. `return detectorspb.DetectorType_AzureBatch` →
# `azurebatch`). Necessary because the package directory often differs
# from the enum name (azure_batch vs AzureBatch, npmtokenv2 vs NpmToken,
# close vs closecrm, etc.).
# - Version comes from the directory suffix only (`/v<n>`). Detectors that
# encode the version in the dir name (e.g. `npmtokenv2`) are emitted
# without a version suffix; trufflehog then matches all versions of that
# proto type — wider scope but correct.
#
# "New detector" detection compares pkg/engine/defaults/defaults.go imports
# between the two refs. A detector imported at HEAD but not at BASE is new.
#
# Modes:
# (none) List all changed detectors at HEAD, one per line, in
# <name>[.v<n>] form.
# --pr-csv Same set as default mode, comma-joined.
# --main-csv Changed detectors that also exist at BASE (excludes new),
# comma-joined. Use as --include-detectors for the main build.
# --new-only Just the new detectors (in HEAD but not BASE), one per line.
#
# Env:
# BASE_REF default origin/main
# HEAD_REF default HEAD
set -euo pipefail
MODE="${1:-list}"
BASE_REF="${BASE_REF:-origin/main}"
HEAD_REF="${HEAD_REF:-HEAD}"
REPO_ROOT="$(git rev-parse --show-toplevel)"
cd "$REPO_ROOT"
# Resolve BASE to a concrete commit. Workflow already runs `git fetch origin
# main`; locally that may not be true, so we fall back to `main` if the
# remote-tracking ref is missing.
if ! git rev-parse --verify "$BASE_REF" >/dev/null 2>&1; then
if git rev-parse --verify main >/dev/null 2>&1; then
BASE_REF=main
else
echo "error: cannot resolve BASE_REF=$BASE_REF and no local 'main'" >&2
exit 1
fi
fi
MERGE_BASE=$(git merge-base "$BASE_REF" "$HEAD_REF")
# Step 1 — changed detector dirs (relative to repo root).
# Pattern: pkg/detectors/<name>(/v<n>)?/<file>.go, excludes _test.go and
# files inside common/, custom_detectors/.
mapfile -t CHANGED_DIRS < <(
git diff --name-only "$MERGE_BASE...$HEAD_REF" -- 'pkg/detectors/**/*.go' \
| grep -Ev '_test\.go$' \
| grep -Ev '^pkg/detectors/(common|custom_detectors)/' \
| sed -E 's|^(pkg/detectors/[^/]+(/v[0-9]+)?)/[^/]+\.go$|\1|' \
| sort -u
)
# Step 2 — defaults.go imports at each ref. Each line has form
# "github.com/trufflesecurity/trufflehog/v3/pkg/detectors/<name>(/v<n>)?"
# We extract just the <name>(/v<n>)? portion to use as the dir identifier.
parse_defaults_imports() {
local ref="$1"
git show "$ref:pkg/engine/defaults/defaults.go" 2>/dev/null \
| grep -oE '"github\.com/trufflesecurity/trufflehog/v3/pkg/detectors/[^"]+"' \
| sed -E 's|.*/pkg/detectors/||; s|"$||' \
| sort -u
}
mapfile -t HEAD_IMPORTS < <(parse_defaults_imports "$HEAD_REF")
mapfile -t BASE_IMPORTS < <(parse_defaults_imports "$MERGE_BASE")
# Set difference: detectors imported at HEAD but not at BASE. The dir
# identifier (e.g. "github/v2", "stripe") matches the form we extracted in
# step 1, so we can intersect directly without re-mapping.
NEW_DIRS_FILE=$(mktemp)
trap 'rm -f "$NEW_DIRS_FILE"' EXIT
comm -23 \
<(printf '%s\n' "${HEAD_IMPORTS[@]+"${HEAD_IMPORTS[@]}"}") \
<(printf '%s\n' "${BASE_IMPORTS[@]+"${BASE_IMPORTS[@]}"}") \
> "$NEW_DIRS_FILE"
is_new_detector() {
grep -qxF "$1" "$NEW_DIRS_FILE"
}
# Step 2b — skip detectors whose diff doesn't touch regex patterns or Keywords.
# Corpora results only change when the matching logic changes; verification,
# redaction, or structural changes don't affect match counts.
has_pattern_change() {
local dir="$1"
# Fast path: regex or Keywords() signature on a changed line.
git diff "$MERGE_BASE...$HEAD_REF" -- "$dir"/*.go 2>/dev/null \
| grep -qE '^[+-][^+-].*(regexp\.|MustCompile|Keywords)' && return 0
# Slow path: compare the Keywords() function body between refs to catch
# changes to the return value (e.g. []string{"old"} → []string{"new"})
# where the changed lines don't mention "Keywords" themselves.
local file
while IFS= read -r file; do
[[ "$file" == *_test.go ]] && continue
local head_body base_body
head_body=$(git show "$HEAD_REF:$file" 2>/dev/null \
| awk '/func[[:space:]].*Keywords\(\)[[:space:]]*\[\]string/,/^[[:space:]]*\}/' \
| tail -n +2)
base_body=$(git show "$MERGE_BASE:$file" 2>/dev/null \
| awk '/func[[:space:]].*Keywords\(\)[[:space:]]*\[\]string/,/^[[:space:]]*\}/' \
| tail -n +2)
[[ "$head_body" != "$base_body" ]] && return 0
done < <(git diff --name-only "$MERGE_BASE...$HEAD_REF" -- "$dir"/*.go 2>/dev/null)
return 1
}
# Step 3 — for a dir, derive `<protoname>[.v<n>]`.
detector_id_for_dir() {
local dir="$1"
local version=""
if [[ "$dir" =~ ^pkg/detectors/[^/]+/v([0-9]+)$ ]]; then
version=".v${BASH_REMATCH[1]}"
fi
# Extract proto enum name. Multiple matches are possible (a detector may
# also reference related types in helpers); the Type() return is by far
# the most common, so the modal value wins.
local proto
proto=$(
grep -E 'return[[:space:]]+\S*DetectorType_[A-Za-z0-9]+' "$dir"/*.go 2>/dev/null \
| grep -v '_test\.go' \
| grep -oE 'DetectorType_[A-Za-z0-9]+' \
| sort | uniq -c | sort -rn \
| head -1 \
| awk '{print $2}' \
| sed 's/^DetectorType_//' \
| tr '[:upper:]' '[:lower:]'
)
if [[ -z "$proto" ]]; then
return 1
fi
echo "${proto}${version}"
}
# Step 4 — emit per mode.
emit_list() {
local dir id
for dir in "${CHANGED_DIRS[@]:-}"; do
[[ -z "$dir" ]] && continue
has_pattern_change "$dir" || continue
if id=$(detector_id_for_dir "$dir"); then
echo "$id"
else
echo "warning: could not resolve detector id for $dir" >&2
fi
done | sort -u
}
emit_main_list() {
local dir id
for dir in "${CHANGED_DIRS[@]:-}"; do
[[ -z "$dir" ]] && continue
has_pattern_change "$dir" || continue
# Strip `pkg/detectors/` prefix to get the import-path form, then
# check against the new-detector set.
local import_form="${dir#pkg/detectors/}"
if is_new_detector "$import_form"; then
continue
fi
if id=$(detector_id_for_dir "$dir"); then
echo "$id"
fi
done | sort -u
}
emit_new_list() {
local dir id
for dir in "${CHANGED_DIRS[@]:-}"; do
[[ -z "$dir" ]] && continue
has_pattern_change "$dir" || continue
local import_form="${dir#pkg/detectors/}"
if ! is_new_detector "$import_form"; then
continue
fi
if id=$(detector_id_for_dir "$dir"); then
echo "$id"
fi
done | sort -u
}
case "$MODE" in
list) emit_list ;;
--pr-csv) emit_list | paste -sd, - ;;
--main-csv) emit_main_list | paste -sd, - ;;
--new-only) emit_new_list ;;
*) echo "Usage: $0 [--pr-csv|--main-csv|--new-only]" >&2; exit 2 ;;
esac