Agent Building skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

agent-observability-eval-bootstrap

Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit Python SDK code or a framework-agnostic JSON spec instead. Use when user says "bootstrap evaluators", "generate evaluators", "create evals from traces", "eval bootstrap", "write evaluators", "build eval suite", "publish evaluators", or wants to generate BaseEvaluator/LLMJudge code or online judge configs from production LLM trace data. Works with ml_app and optional RCA report or failure hypothesis.

142 repo starsObserved in 2 repos
Agent Building

agent-observability-eval-pipeline

End-to-end Agent Observability pipeline for an instrumented ml_app — classify production traces, root-cause failures, bootstrap evaluators, then (optionally) sample + publish a dataset, generate + run an experiment, and analyze results. Six narrated phases with a standardized banner and a "continue" checkpoint between each. Pure orchestration over the agent-observability sub-skills (`agent-observability-session-classify`, `agent-observability-trace-rca`, `agent-observability-eval-bootstrap`, `agent-observability-experiment-py-bootstrap`, `agent-observability-experiment-analyzer`). Use when user says "run the eval pipeline", "go from traces to evals", "bootstrap evals end to end", "classify then RCA then bootstrap", "build an eval set from scratch", "onboard me to datasets and experiments", "walk me through experiments", "I have an ml_app, now what", "Agent Observability onboarding", "guided experiment setup", "from traces to experiments", or wants a deterministic, narrated tour from production data through evaluators, datasets, and experiments. Stop early with `--stop-after <phase>` to short-circuit at evaluators or dataset, or resume mid-flow with `--start-at <phase>`.

142 repo starsObserved in 2 repos
Agent Building

agent-observability-experiment-analyzer

Analyze LLM experiment results. Handles single or comparative experiments, exploratory or Q&A modes. Use when user says "analyze experiment", "compare experiments", "analyze against baseline", or provides one or two experiment IDs for analysis.

142 repo starsObserved in 2 repos
Agent Building

agent-observability-session-classify

Classify whether user intent was satisfied in a Datadog Agent Observability trace or session. Three modes: (1) session_id — classify a single CMD+I assistant session with RUM; (2) trace_id — classify a single Agent Observability trace without RUM; (3) ml_app — sample and classify multiple sessions or traces from a given LLM app. Output is compact by default (verdict + one-sentence reason). Use when evaluating satisfaction, classifying sessions/traces, labeling data, or generating signal for agent-observability-eval-pipeline or agent-observability-trace-rca.

142 repo starsObserved in 2 repos
Agent Building

system-prompt-structure

Anatomy of effective system prompts — role, context, constraints, format.

142 repo starsObserved in 2 repos
Agent Building

openai-ai

Manage OpenAI files, assistants, vector stores, batches, fine-tuning jobs, and model resources via the OpenAI API. Use this skill when users want to create or manage assistants, upload files, run batch jobs, fine-tune models, generate images or audio, and work with the Assistants API via OpenAI.

142 repo starsObserved in 1 repos
Agent Building

soul-md-creator

Create or improve SOUL.md files for OpenClaw agents. Use when the user wants to design an agent personality, rewrite an existing soul, align SOUL.md with IDENTITY.md, or prepare a soul for publishing on souls.directory with optional frontmatter.

142 repo starsObserved in 1 repos
Agent Building

test-agent-hooks

Exercise a real agent hook turn that drives Zedra Running, WaitingApproval, and Completed state.

142 repo starsObserved in 1 repos
Agent Building

mem0

Mem0 Platform SDK for adding persistent memory to AI applications. TRIGGER when: user mentions "mem0", "MemoryClient", "memory layer", "remember user preferences", "persistent context", "personalization", or needs to add long-term memory to chatbots, agents, or AI apps. Covers Python SDK (mem0ai), TypeScript SDK (mem0ai), and framework integrations (LangChain, CrewAI, OpenAI Agents SDK, Pipecat, LlamaIndex, AutoGen, LangGraph). Also covers the open-source self-hosted Memory class. This is the DEFAULT mem0 skill for ambiguous queries. DO NOT TRIGGER when: user asks about CLI commands, terminal usage, or shell scripts (use mem0-cli), or Vercel AI SDK / @mem0/vercel-ai-provider / createMem0 (use mem0-vercel-ai-sdk).

141 repo starsObserved in 14 repos
Agent Building