Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

siggraph-artifact-evaluation

Use when pursuing the Graphics Replicability Stamp (GRSI) or Code Replicability in Computer Graphics (CRCG) recognition for an accepted SIGGRAPH / TOG paper, covering how graphics replicability differs from ACM artifact badging, what volunteers actually run, deterministic result reproduction, Software Heritage archiving, and the separate post-acceptance timing.

805 repo starsObserved in 2 repos
Testing & Quality

siggraph-reproducibility

Use when building the reproducibility story for a SIGGRAPH / TOG paper, covering deterministic result regeneration, scene/mesh/weight provenance, hardware and timing disclosure, floating-point and GPU non-determinism, and a code/data release that a reader or a Graphics Replicability Stamp volunteer can actually run.

805 repo starsObserved in 2 repos
Testing & Quality

sigmetrics-artifact-evaluation

Use when packaging an ACM SIGMETRICS artifact for the ACM Artifact Review and Badging scheme (Artifacts Available, Evaluated Functional and Reusable, Results Reproduced), covering what performance-evaluation evaluators check first (does the simulation regenerate the figures and match the analysis?), DOI-issuing archives, evaluator-proof documentation, and confirming whether an artifact track runs this cycle.

805 repo starsObserved in 2 repos
Testing & Quality

sigmetrics-reproducibility

Use when strengthening ACM SIGMETRICS reproducibility, covering proofs and their assumptions as reproducible artifacts, seeded simulators whose figures regenerate and match the analysis, measurement/trace provenance, claim-to-evidence mapping, honest degrees of reproducibility, and consistency between what the paper proves/measures and what the artifact contains.

805 repo starsObserved in 2 repos
Testing & Quality

sigmod-experiments

Use when designing or auditing the evaluation of a SIGMOD paper, covering workload realism and standard benchmark usage, baseline tuning fairness, scalability and tail-latency methodology, ablations that isolate the mechanism, and the setup disclosure a data-systems PC demands before trusting any speedup.

805 repo starsObserved in 2 repos
Testing & Quality

socc-artifact-evaluation

Use when packaging an ACM SoCC artifact for the ACM Artifact Review and Badging scheme (Artifacts Available, Evaluated Functional and Reusable, Results Reproduced), covering what a cloud-systems evaluator checks first, reproducing tail-latency and cost results on a testbed, DOI-issuing archives, and the fact that whether SoCC runs a dedicated artifact-evaluation track for a given edition must be verified.

805 repo starsObserved in 2 repos
Testing & Quality

socc-experiments

Use when designing or auditing ACM SoCC evaluations, covering real or realistic deployments over simulation, production or representative workloads and traces, tail-latency and cost as first-class metrics, fair tuned baselines, scale and multi-tenancy behavior, reproducible measurement pipelines, and matching evidence to the shape of each cloud claim.

805 repo starsObserved in 2 repos
Testing & Quality

socc-reproducibility

Use when strengthening ACM SoCC reproducibility, covering the testbed and workload description, released code and traces, provenance pinning for measurement studies, reproducing tail-latency and cost (not just the mean), claim-to-evidence mapping, honest degrees of reproducibility, and consistency between what the paper reports and what the artifact regenerates.

805 repo starsObserved in 2 repos
Testing & Quality

soda-reproducibility

Use when hardening the verifiability of a SODA (ACM-SIAM Symposium on Discrete Algorithms) paper, where reproducibility means checkable mathematics — complete proofs in the submitted full version, stable statement-proof correspondence, explicit constants and model assumptions, and certificates for any machine-checked step.

805 repo starsObserved in 2 repos
Testing & Quality

sosp-experiments

Use when designing or auditing the evaluation of a SOSP paper — mapping every claim to an experiment, choosing baselines a systems PC will accept as fair, mixing microbenchmarks with end-to-end and failure runs, reporting tails and overheads honestly, and isolating the mechanism the design credits.

805 repo starsObserved in 2 repos
Testing & Quality