Total 54,172 skills, Data Processing has 2771 skills
Showing 12 of 2771 skills
Use Crawl4AI for web crawling, markdown extraction, and LLM-powered structured extraction through OpenRouter. Use when the user mentions Crawl4AI, unclecode/crawl4ai, wants website data extracted with Crawl4AI, or needs an agent to crawl pages and turn them into structured JSON with OpenRouter-backed models.
TransForm integration. Manage data, records, and automate workflows. Use when the user wants to interact with TransForm data.
This skill should be used when the user asks to "use marimo", "create a marimo notebook", "debug a marimo notebook", "inspect cells", "understand reactive execution", "fix marimo errors", "convert from jupyter to marimo", or works with marimo reactive Python notebooks.
Tealium platform help — enterprise Customer Data Platform (CDP), Real-Time CDP, Composable CDP, iQ Tag Management, EventStream, AudienceStream, 1300+ connectors, identity resolution, consent management, V3 API. Use when setting up Tealium CDP, tags not firing or loading slowly, connectors failing or data not flowing to destinations, visitor profiles not merging across channels, identity resolution producing duplicate profiles, choosing between Tealium and Segment or mParticle, configuring EventStream or AudienceStream, or working with the Tealium V3 API. Do NOT use for general email marketing (use /sales-email-marketing) or CRM data dedup without Tealium (use /sales-data-hygiene).
Finage integration. Manage data, records, and automate workflows. Use when the user wants to interact with Finage data.
Expertise in generating clean, correct, and efficient Dataform pipeline code for BigQuery ELT. Use this when creating or modifying Dataform pipelines, actions, or source declarations, when Dataform, SQLX, or BigQuery are mentioned in a transformation, when data needs to be ingested from GCS into BigQuery via Dataform, or when setting up a new Dataform project or configuring workflow_settings.yaml.
Data validation using Great Expectations. Expectation suites, checkpoints, and data docs for pipeline monitoring.
Design an end-to-end MotherDuck pipeline. Use when choosing raw, staging, and analytics boundaries, bulk ingestion paths, transformation sequencing, publication targets, or whether DuckLake is actually required.
Analyze unit economics for PE targets — ARR cohorts, LTV/CAC, net retention, payback periods, revenue quality, and margin waterfall. Essential for software/SaaS, recurring revenue, and subscription businesses. Use when evaluating revenue quality, building a cohort analysis, or assessing customer economics. Triggers on "unit economics", "cohort analysis", "ARR analysis", "LTV CAC", "net retention", "revenue quality", or "customer economics".
Produces a CPA-ready year-end COGS schedule for Amazon sellers from Inventory Valuation Report, Settlement reports, and supplier POs. Catches the ending inventory math errors that overstate taxable income by 8-15%. Use when a user asks about year-end taxes, COGS, inventory valuation, or "tax pack for my accountant". Trigger phrases: "year-end taxes", "COGS schedule", "inventory valuation", "FIFO LIFO Amazon", "tax pack for CPA". Works with zero tools.
Access Human Metabolome Database (220K+ metabolites). Search by name/ID/structure, retrieve chemical properties, biomarker data, NMR/MS spectra, pathways, for metabolomics and identification.
Work with Data Commons, a platform providing programmatic access to public statistical data from global sources. Use this skill when working with demographic data, economic indicators, health statistics, environmental data, or any public datasets available through Data Commons. Applicable for querying population statistics, GDP figures, unemployment rates, disease prevalence, geographic entity resolution, and exploring relationships between statistical entities.