Detects repeated automated code-review/fix loops and forces a deeper invariant-ledger stabilization pass before more review requests. Use during Grace issue/PR orchestration when Codex or another reviewer repeatedly finds substantive adjacent defects, especially in storage, actors, persistence, retries, concurrency, authorization, public contracts, or other high-risk code.
Use when inspecting a Java program that is already being debugged — read local variables, walk the call stack, list threads, evaluate expressions, step in/over/out, continue execution, or set / remove breakpoints in an active Java debug session. NOT for starting, launching, or stopping a debug session — use `java-launch-troubleshooting` for that.
Fix code quality issues identified in a code quality review stored in agent_artefacts/code_quality/<topic>/. Systematically addresses issues found by the code-quality-review-all skill for ANY code quality topic, with validation and testing at each step. Use when user asks to fix issues from a code quality review, or asks to fix issues from agent_artefacts/code_quality/<topic>.
Review all evaluations in the repository against a single code quality standard. Checks ALL evals against ONE standard for periodic quality reviews. Use when user asks to review/audit/check all evaluations for a specific topic or standard. Do NOT use for reviewing a single eval (use eval-quality-workflow instead) or for test coverage (use ensure-test-coverage instead).
Ensure test coverage for a single evaluation - both reviewing existing tests and creating missing ones. Analyzes testable components, checks tests against repository conventions, reports coverage gaps, and creates or improves tests. Use when user asks to check/review/create/add/ensure tests for an eval. Use whenever you are asked to review an evaluation that contains tests, or whenever you need to write a suite of tests. Do NOT use for fixing a specific failing CI test (use ci-maintenance-workflow instead).
Fix or review a single evaluation against all EVALUATION_CHECKLIST.md standards. Use "fix" mode to refactor an eval into compliance, or "review" mode to assess compliance without making changes. Use when user asks to fix, review, or check an evaluation's quality. Trigger when the user asks you to run the "Fix An Evaluation" or "Review An Evaluation" workflow. Do NOT use for reviewing ALL evals against a single code quality standard (use code-quality-review-all instead).
Review a single evaluation's validity — whether its claims hold up, whether its name is accurate, whether samples can be both succeeded and failed at, and whether scoring measures ground truth. Use when user asks to check validity of an eval, or as part of the Master Checklist workflow. Do NOT use for code quality or test coverage (use eval-quality-workflow or ensure-test-coverage instead).
View and analyse Inspect evaluation log files using the Python API. Trigger whenever you need to look at a .eval file yourself without using pre-written scripts.
Analyse the SUB/WAVE radio station's unified event log (state/logs/events-*.jsonl) and give the operator a diagnostic report on how the station is behaving — Navidrome/ Subsonic API call patterns, the DJ picker's behaviour and music-library coverage, and runtime health anomalies. Use this skill whenever the user wants to understand or get feedback on what the radio is doing under the hood: how Navidrome calls are being made and how often, why the picker keeps choosing certain tracks or artists, whether the library pool is too narrow, whether traces are failing or running slow, or asks things like "analyse the subwave logs", "check the radio's behaviour", "what is the picker doing", "why does it keep playing the same artists", "is the station healthy", "how are the navidrome calls looking", "give me feedback on the event log". Trigger this skill even if the user does not name the log file or the script — any request to diagnose, audit, review, or get feedback on SUB/WAVE's runtime behaviour from its logs belongs here.