Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

aod-failure-archaeology

Use before touching decay half-life, prior/global-prior calculation, sensor-health repairs, sleep-presence detection, recorder/DB write volume, adjacent-areas, timezone/DST-sensitive datetime code, or the config-flow advanced-options gating — to check whether the bug you're about to fix (or the fix you're about to write) already happened before. Load this when a symptom rhymes with "custom value silently reverted to a default", "repairs fire spuriously / return after Ignore", "prior pinned near 0.99 or 0.01", "TypeError offset-naive and offset-aware", or "recorder database growing too fast".

309 repo starsObserved in 1 repos
Testing & Quality

aod-learning-accuracy-campaign

Executable, decision-gated campaign for making learned priors and likelihoods (global prior, 168 time-priors, per-sensor P(E|H) from correlation analysis, decay half-life resolution) trustworthy on real homes. Load this when a user reports "prior stuck at 0.99/0.01", "occupancy probability doesn't match reality", "false transitions", "decay clears too fast/slow", "correlation/likelihood looks wrong", or when asked to touch data/prior.py, data/analysis.py's PriorAnalyzer, db/correlation.py, or data/decay.py half-life resolution. This is the flagship operation order for the project's hardest live problem — treat it as a runbook, not background reading.

309 repo starsObserved in 1 repos
Testing & Quality

aod-proof-and-analysis-toolkit

Use when you are about to touch, review, or debug any probability/statistics math in Area Occupancy Detection (priors, likelihoods, decay half-lives, logit-space boosts, correlation confidence, timezone/interval bucketing) and need to hand-verify the numbers BEFORE trusting a PR, a bug report, or your own patch. Trigger on tasks like "does this prior look right", "is this decay curve correct", "why did the prior pin at 0.99", "check this ratio/denominator", "audit this new coupling between areas for feedback loops", or any request to prove a calculation from first principles rather than just running the test suite.

309 repo starsObserved in 1 repos
Testing & Quality

aod-validation-and-qa

Use when writing, reviewing, or deciding what evidence is required for ANY change to Area Occupancy Detection — before opening a PR, when asked "is this tested enough?", when adding a test for math/behavior/config/DB code, when coverage or CI is failing, or when a reviewer/CodeRabbit flags a test as rigged, redundant, or encoding a bug. Covers the evidence bar per change class, conftest.py fixture anatomy, the golden test-file map, and lint/coverage gates.

309 repo starsObserved in 1 repos
Testing & Quality

code-refactor-review

Reviews code changes for reuse, composition, codebase consistency, and slop. Use when asked to review PRs/diffs, check code reuse, composition, cleanliness, or whether code fits the codebase.

309 repo starsObserved in 1 repos
Testing & Quality

fault-injection-loop

Closed-loop fault injection and attribution accuracy benchmark. Draws from a prioritized pool of (fault_type, rank, iter, nodes) experiments and submits them 2 at a time via sbatch — waiting for each pair to finish before submitting the next — to bound filesystem load. GPU-related faults are front-loaded in the pool. After all jobs complete, runs /log-analysis and /fr-analysis on every experiment, scores attribution vs. ground truth, aggregates gaps, and iterates on attribution modules to close them.

308 repo starsObserved in 1 repos
Testing & Quality

fr-analysis

Analyze PyTorch NCCL flight-recorder (FR) dumps to identify collective operation hangs and isolate the responsible ranks using CollectiveAnalyzer. Use when a distributed training job hangs due to an NCCL collective timeout and FR dump files are available. Detects the wavefront process group where collectives diverge and returns the root-cause suspect ranks.

308 repo starsObserved in 1 repos
Testing & Quality

nvrx-attr

Orchestration layer over nvidia_resiliency_ext attribution modules. Provides log-analysis, fr-analysis, and a Megatron-LM-oriented fault-injection feedback loop for benchmarking attribution quality on SLURM workloads.

308 repo starsObserved in 1 repos
Testing & Quality

rspec-testing

Bike Index's RSpec testing conventions — how to structure specs with `context` and `let`, what kinds of tests to write, and what to avoid (mocks, controller specs, testing private methods). Trigger when writing or modifying any `*_spec.rb` file, adding test coverage for new code, refactoring tests, or designing the test layout for a new feature. Includes Good/Bad examples of the project's preferred style.

308 repo starsObserved in 1 repos
Testing & Quality

qt-qml-profiler

Use when the user is investigating QML / Qt Quick performance — both vague complaints ("the UI feels laggy", "this is slow", "frames are dropping", "the app stutters") and explicit asks to profile, find hotspots, or optimize bindings, signals, or rendering. Runs qmlprofiler on a 2D QML application, parses the .qtd trace, and analyzes hotspots against the source with frame-time, memory, and pixmap-cache summaries. Does NOT cover Qt Quick 3D.

307 repo starsObserved in 7 repos
Testing & Quality