Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

micro-reproducibility

Use when hardening a MICRO paper's results for re-derivation — pinning simulator commits and configs, recording workload trace provenance and SimPoint recipes, versioning power/area models, capturing RTL toolchain state, and writing the run manifests that later survive MICRO's post-acceptance artifact evaluation.

805 repo starsObserved in 2 repos
Testing & Quality

mlsys-artifact-evaluation

Use when packaging an accepted MLSys paper's code, configs, and measurement scripts for the venue's post-acceptance artifact evaluation, targeting the Availability, Functional, and Reproducible badges, writing the Artifact Appendix, handling hardware that AE reviewers cannot access, and answering anonymous evaluator questions.

805 repo starsObserved in 2 repos
Testing & Quality

mlsys-experiments

Use when designing or auditing the evaluation of an MLSys paper, selecting representative workloads and hardware, tuning baselines symmetrically, reporting throughput, latency tails, memory, cost, and quality together, structuring ablations that attribute gains to mechanisms, and building scaling and sensitivity evidence reviewers trust.

805 repo starsObserved in 2 repos
Testing & Quality

mlsys-reproducibility

Use when hardening the reproducibility of MLSys performance claims, pinning the full system layer from driver to interconnect, separating ML randomness from systems noise, choosing repetition counts and variance reporting for throughput and latency numbers, and disclosing hardware, workloads, and cost so strangers can re-measure results.

805 repo starsObserved in 2 repos
Testing & Quality

mobisys-experiments

Use when designing or auditing the evaluation of a MobiSys submission — building real-device testbeds, instrumenting energy and thermal behavior, measuring latency and frame-rate tails, bounding memory footprint, choosing tuned system baselines, and running deployments or user studies, so systems reviewers see where the system wins and breaks on the device.

805 repo starsObserved in 2 repos
Testing & Quality

nsdi-experiments

Use when designing or auditing the evaluation of an NSDI submission — choosing traces, testbeds, and deployment evidence, sizing scale and failure-injection experiments, picking baselines that fight back, and reporting tail behavior so networked-systems reviewers can see where the design wins, loses, and breaks.

805 repo starsObserved in 2 repos
Testing & Quality

osdi-experiments

Use when designing or auditing the evaluation of an OSDI submission — choosing mature baselines and realistic workloads, structuring the section around research questions, measuring scalability and tail behavior, quantifying the design's costs, and fitting the evidence into the 12-page reviewed body.

805 repo starsObserved in 2 repos
Testing & Quality

percom-artifact-evaluation

Use when packaging an IEEE PerCom sensing artifact and dataset for reproducibility and any badging (IEEE Open Research Objects / Results Reproduced, IEEE DataPort or Zenodo deposit), covering what a ubicomp evaluator checks first for human-subjects sensing data, cross-subject reproduction, de-identification, and honest degrees of reproducibility.

805 repo starsObserved in 2 repos
Testing & Quality

pldi-experiments

Use when designing or auditing a PLDI evaluation — choosing defensible benchmark suites and baseline compiler configurations, measuring runtime, compile time, and memory with warmup and variance discipline, running ablations that isolate the claimed mechanism, and scoping claims to the platforms measured.

805 repo starsObserved in 2 repos
Testing & Quality

pldi-reproducibility

Use when hardening a PLDI paper's measurements against the SIGPLAN Empirical Evaluation Guidelines — warmup and steady-state discipline, variance and confidence reporting, principled benchmark choice, pinned toolchains, cross-platform validity, and a measurement log that survives artifact evaluation.

805 repo starsObserved in 2 repos
Testing & Quality