Loading...
Loading...
Found 1,893 Skills
Instrument Python LLM apps, build golden datasets, write eval-based tests, run them, and root-cause failures — covering the full eval-driven development cycle. Make sure to use this skill whenever a user is developing, testing, QA-ing, evaluating, or benchmarking a Python project that calls an LLM, even if they don't say "evals" explicitly. Use for making sure an AI app works correctly, catching regressions after prompt changes, debugging why an agent started behaving differently, or validating output quality before shipping.
Help users create and run AI evaluations. Use when someone is building evals for LLM products, measuring model quality, creating test cases, designing rubrics, or trying to systematically measure AI output quality.
LLM observability platform for tracing, evaluation, prompt management, and cost tracking. Use when setting up Langfuse, monitoring LLM costs, tracking token usage, or implementing prompt versioning.
Build and run evaluators for AI/LLM applications using Phoenix.
Coordinate multi-agent code review with specialized perspectives. Use when conducting code reviews, analyzing PRs, evaluating staged changes, or reviewing specific files. Handles security, performance, quality, and test coverage analysis with confidence scoring and actionable recommendations.
Best practices for scikit-learn machine learning, model development, evaluation, and deployment in Python
Evaluate any address for home buyers and renters. Get nearby schools, transit, grocery stores, parks, restaurants, and walkability using Camino AI's location intelligence.
System architecture and technical design specialist. 🚨 TIER 2 SKILL - ON-DEMAND ACTIVATION 🚨 Use when user requests involve: - System architecture design and planning - Technical specifications and ADRs - Technology evaluation and selection - Scalability and performance planning - Integration architecture and API design - English: "design system", "architecture", "ADR", "tech stack", "scalability" - Swedish: "arkitektur", "systemdesign", "teknikval", "skalbarhet" Architecture Specialist (British female voice) provides: - System design and architecture patterns - Architecture Decision Records (ADRs) - Technology evaluation and trade-off analysis - Cloud and microservices architecture - Integration patterns and API design User confirmation optional but recommended for major architectural decisions.
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
Investment Analysis: Generate an in-depth investment analysis report. We do not conduct traditional investment analysis—the core judgment is whether the project is an "Order-Creating Machine". Activate this when the user says "investment report", "investment analysis", "analyze this project", "write an investment report", "investment report", "invest analysis", or provides entrepreneur conversation records for investment evaluation. Also activate when the user pastes or references meeting notes, pitch decks, or founder interviews and requests analysis.
Produces a CPA-ready year-end COGS schedule for Amazon sellers from Inventory Valuation Report, Settlement reports, and supplier POs. Catches the ending inventory math errors that overstate taxable income by 8-15%. Use when a user asks about year-end taxes, COGS, inventory valuation, or "tax pack for my accountant". Trigger phrases: "year-end taxes", "COGS schedule", "inventory valuation", "FIFO LIFO Amazon", "tax pack for CPA". Works with zero tools.
Analyze dividend investment opportunities, evaluate dividend safety, growth potential and yield rate. Use this when users inquire about dividends, dividend investment or dividend yield. Supports quick screening, in-depth analysis and portfolio optimization.