Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

cvpr-experiments

Use when designing or auditing the experimental program of a CVPR paper, covering benchmark and baseline selection under matched-compute fairness, the ablation study reviewers treat as mandatory, qualitative and failure-case evidence, efficiency metrics tied to the Compute Reporting Form, and generalization tests beyond a single dataset.

805 repo starsObserved in 2 repos
Testing & Quality

dac-experiments

Use when designing or auditing the empirical evaluation of an ACM/IEEE Design Automation Conference (DAC) Research Manuscript, covering standard EDA benchmark suites (ISPD, EPFL, ISCAS/ITC, TAU, CircuitNet, OpenROAD flows), fair state-of-the-art baselines, QoR/PPA reporting with runtime, per-benchmark honesty, ablations that isolate the mechanism, and contamination-aware ML-for-EDA evaluation.

805 repo starsObserved in 2 repos
Testing & Quality

ectj-identification-strategy

Use when stress-testing identification, assumptions, asymptotics, regularity conditions, and proofs in a The Econometrics Journal (EctJ) submission, including proof placement under RES printed-appendix rules and pairing every asymptotic claim with finite-sample evidence referees can audit.

805 repo starsObserved in 2 repos
Testing & Quality

edbt-experiments

Use when designing or auditing EDBT empirical evaluations for database-systems work, covering real workloads and datasets, fair and tuned baselines, honest measurement across realistic scales, reproducible harnesses, and the higher bar of the Experiments & Analysis paper where the measurement itself is the contribution.

805 repo starsObserved in 2 repos
Testing & Quality

edbt-reproducibility

Use when strengthening EDBT reproducibility for a database-systems paper, covering a runnable artifact, pinned environments and workloads, dataset and query-log provenance, claim-to-evidence mapping, honest degrees of reproducibility, and consistency between the paper and the package for the open-access OpenProceedings record.

805 repo starsObserved in 2 repos
Testing & Quality

eer-robustness

Use when a European Economic Review (EER) result must be shown to survive specification, sample, measurement, and inference changes — the robustness battery referees demand. Builds the stress tests and organizes them; it does not establish the core identification or write the prose.

805 repo starsObserved in 2 repos
Testing & Quality

eurosys-artifact-evaluation

Use when preparing a EuroSys artifact for the sysartifacts-run evaluation — choosing among the Available, Functional, and Reproduced badges, timing the post-notification artifact submission, building for an evaluator on foreign hardware, and aiming at the Gilles Muller Best Artifact Award rather than a minimal pass.

805 repo starsObserved in 2 repos
Testing & Quality

fast-experiments

Use when designing or auditing a USENIX FAST storage evaluation, covering real devices and firmware, device-state control (aging, preconditioning, fill, TRIM), standard workloads and traces (SNIA IOTTA, YCSB, filebench, fio), write amplification, tail latency, endurance and wear, crash-consistency testing, fair baselines, and matching the metric to the shape of each storage claim.

805 repo starsObserved in 2 repos
Testing & Quality

focs-reproducibility

Use when hardening a FOCS (IEEE Symposium on Foundations of Computer Science) paper's checkability — the theory analogue of reproducibility — via hypothesis ledgers, single-source theorem statements that cannot drift between submission and arXiv versions, audits of imported theorems, and certificates for machine-checked steps.

805 repo starsObserved in 2 repos
Testing & Quality

icassp-experiments

Use when designing or auditing ICASSP experiments across signal-processing modalities — matching the metric to the task law (WER, SI-SDR, PESQ/STOI, EER/minDCF, PSNR/SSIM, BER, RMSE), anchoring baselines to current strong methods and standard corpora, sweeping the operating condition, and reporting spread over runs within the four-page limit.

805 repo starsObserved in 2 repos
Testing & Quality