Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

stoc-reproducibility

Use when hardening a STOC (ACM Symposium on Theory of Computing) paper so its results can be independently checked — proof completeness across the extended-abstract/full-version split, single-source builds that prevent statement drift between the two documents, and determinism for any computation a claim relies on.

805 repo starsObserved in 2 repos
Testing & Quality

tacas-experiments

Use when designing or auditing a TACAS (ETAPS) evaluation, covering shared verification benchmarks (SV-COMP-style task sets), fair baseline configuration and equal time budgets, honest wall-clock/scalability reporting on stated hardware, soundness checking of results, reproducibility on the clean artifact VM, and how a TACAS tool-paper evaluation differs from a SV-COMP competition entry.

805 repo starsObserved in 2 repos
Testing & Quality

tacas-reproducibility

Use when strengthening TACAS (ETAPS) reproducibility, covering the clean evaluation-VM packaging that the artifact process assumes, pinned dependencies and offline execution, a claim-to-script mapping so every benchmark number regenerates, honest degrees of reproducibility, consistency between the paper and the artifact, and the category difference between a mandatory tool-paper artifact and a voluntary research-paper artifact.

805 repo starsObserved in 2 repos
Testing & Quality

uist-experiments

Use when designing or auditing the evaluation of a UIST paper — choosing among technical benchmarks, controlled comparisons, usability walkthroughs, expert sessions, and demonstration applications, matching evaluation shape to the systems claim, and avoiding the ritual study that proves nothing the paper asserts.

805 repo starsObserved in 2 repos
Testing & Quality

vldb-experiments

Use when designing or auditing the evaluation of a VLDB paper, covering workload and dataset realism at scale, competitor tuning fairness, scalability curves versus single points, tail-latency and throughput reporting, ablations that isolate the mechanism, and the loss-case disclosure PVLDB reviewers look for first.

805 repo starsObserved in 2 repos
Testing & Quality

wacv-experiments

Use when designing or auditing WACV experiments, covering Applications-track systems evidence (latency, power, robustness under real constraints) versus Algorithms-track matched-baseline novelty, comparative assessment under the deployed condition, uncertainty over seeds and sessions, ablations, and evidence that survives the two-round review.

805 repo starsObserved in 2 repos
Testing & Quality

wsdm-experiments

Use when designing or auditing the evaluation of a WSDM paper - offline ranking and recommendation metrics with bias controls, temporal-split protocols for interaction logs, baseline selection from recent WSDM/SIGIR/KDD editions, ablations that isolate the mechanism, efficiency reporting, and online-evidence framing.

805 repo starsObserved in 2 repos
Testing & Quality

atc-reproducibility

Use when building the reproducibility story for an ATC (ACM SIGOPS Annual Technical Conference, formerly USENIX ATC) systems paper — pinning testbed and software environments, providing a turnkey path from the artifact to the headline numbers, and preparing an anonymized-but-runnable review package ahead of the Available/Functional/Reproduced badges.

805 repo starsObserved in 1 repos
Testing & Quality

cav-reproducibility

Use when strengthening CAV (Computer Aided Verification) reproducibility, covering benchmark provenance (SV-COMP/SMT-COMP/HWMCC/VNN-COMP set revisions), pinned tool and baseline versions, resource limits and hardware, seeds for randomized/portfolio solvers, checkable proof witnesses/certificates for soundness claims, and consistency between the paper's tables and the artifact.

805 repo starsObserved in 1 repos
Testing & Quality

checking-model-compliance

Checks a Simulink model against a compliance standard (MISRA, MAB, JMAAB, ISO 26262, ISO 25119, DO-178C, DO-254, IEC 61508, IEC 62304, EN 50128, CERT C/CWE, AUTOSAR) using Model Advisor, then summarizes findings and suggests fixes. Use when the user asks whether their model is compliant, wants to run standard checks, or needs a compliance report.

805 repo starsObserved in 1 repos
Testing & Quality