Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

podc-reproducibility

Use when making an ACM PODC paper's result independently checkable — reproducibility for a proofs venue, not an artifact venue. Covers a self-contained proof appendix, an explicit and checkable model/assumption box, honest handling of any optional simulation, and keeping the full version (with proofs) in sync with the 10-page camera-ready.

805 repo starsObserved in 2 repos
Testing & Quality

pods-reproducibility

Use when strengthening the verifiability of an ACM PODS paper — a complete claim-to-proof map, self-contained proofs in the at-submission appendix (no external appendices), correctly stated assumptions, honest scope and open cases, the full-version-on-arXiv norm, and consistency between what the paper claims and what the proofs actually establish.

805 repo starsObserved in 2 repos
Testing & Quality

ppopp-experiments

Use when designing or auditing a PPoPP paper's evaluation, covering the twin bar of concurrency correctness and measured scalability — speedup curves, strong vs weak scaling, core/thread sweeps, NUMA and GPU effects, contention microbenchmarks plus real workloads, variance and measurement hygiene, and honest strong baselines.

805 repo starsObserved in 2 repos
Testing & Quality

ppopp-reproducibility

Use when making a PPoPP paper's parallel-performance results reproducible, covering the hardware and topology description reviewers re-run, thread pinning and NUMA control, seeds and warm-up, compiler/driver/flag provenance, and building an environment that reproduces the paper's scaling trend on different machines.

805 repo starsObserved in 2 repos
Testing & Quality

prai-reproducibility

在为《模式识别与人工智能》(Pattern Recognition and Artificial Intelligence, PR&AI) 准备可复现性与实验可重复材料时调用。覆盖固定随机种子、锁定软硬件环境与依赖版本、固定数据集划分与预处理、预训练权重与大模型版本的记录、复现包(脚本/配置/环境文件)的组织、大数据与权重的稳定托管、以及可复现自查冒烟测试。用于让本刊模式识别/机器学习论文的实验结果能被外审与后续读者重复,降低"结果无法复现"的信任风险;本刊是否设强制复现要求以官网当期须知为准(待核实)。

805 repo starsObserved in 2 repos
Testing & Quality

recsys-reproducibility

Use when strengthening the reproducibility of an ACM RecSys paper or preparing a RecSys Reproducibility Track submission — pinning dataset versions and splits, tuning baselines under an equal budget, reporting seeds and variance, avoiding sampled-metric distortion, and structuring a reproduction study with honest divergence analysis.

805 repo starsObserved in 2 repos
Testing & Quality

rfs-robustness

Use when results may be fragile or when multiple-testing / out-of-sample discipline is the bottleneck for a The Review of Financial Studies (RFS) manuscript. Builds the robustness battery referees will demand; does NOT design identification or write the rebuttal.

805 repo starsObserved in 2 repos
Testing & Quality

rt-simulated-referee

Use to rehearse peer review before submitting — a calibrated Associate Editor desk-screen plus 2–3 distinct-lens referees for the target venue, adversarially verified and synthesized into a referee report, a decision band, and a prioritized, skill-mapped fix list. A rehearsal that predicts the attack surface; it does not replace real review.

805 repo starsObserved in 2 repos
Testing & Quality

sensys-experiments

Use when designing or auditing a SenSys evaluation — energy and low-power measurement with a named instrument, real-testbed and deployment realism, honest sensor ground truth, on-device latency and memory, and same-hardware baselines, so the evidence meets SenSys's built-and-measured bar rather than a simulation or offline-benchmark one.

805 repo starsObserved in 2 repos
Testing & Quality

sigcomm-reproducibility

Use when strengthening the reproducibility evidence of an ACM SIGCOMM paper — topology and testbed ledgers, traffic workload and trace provenance, configuration and version pinning, tail-percentile run counts and variance, legal data-release decisions, and consistency between the paper's claims and the artifact that backs them.

805 repo starsObserved in 2 repos
Testing & Quality