by affaan-m
clickhouse-io is a ClickHouse-focused skill for schema design, analytical SQL, ingestion patterns, and performance tuning. Use it to guide MergeTree choices, partitioning, materialized views, and workload-specific query optimization.
Data Engineering taxonomy generated by the site skill importer.
by affaan-m
clickhouse-io is a ClickHouse-focused skill for schema design, analytical SQL, ingestion patterns, and performance tuning. Use it to guide MergeTree choices, partitioning, materialized views, and workload-specific query optimization.
by ComposioHQ
Snowflake Automation helps agents use Composio MCP to discover Snowflake databases, browse schemas and tables, run SQL, and manage database engineering workflows with role, warehouse, filter, timeout, and safety context.
by ComposioHQ
big-data-cloud-automation helps agents automate Big Data Cloud tasks through Composio Rube MCP by discovering current tool schemas, checking connections, and planning safer execution.
by wshobson
airflow-dag-patterns helps design production-ready Apache Airflow DAGs with stronger task patterns, dependencies, operators, sensors, testing, and deployment guidance for scheduled jobs.
by wshobson
The data-quality-frameworks skill helps teams plan production data validation with dbt tests, Great Expectations, and data contracts. Use it to choose the right checks, map them to a testing pyramid, and guide CI/CD-ready data quality workflows for Data Cleaning and pipeline reliability.
by wshobson
dbt-transformation-patterns helps agents structure dbt projects with staging, intermediate, and marts layers, plus testing, documentation, and incremental model guidance. Use it to plan installs, scaffold new repos, or refactor SQL into cleaner analytics engineering patterns for Database Engineering teams.
by wshobson
spark-optimization is a practical guide to diagnosing slow Apache Spark jobs with partitioning, shuffle, skew, caching, and memory tuning. Use it to install the skill from wshobson/agents, read SKILL.md, and apply evidence-based fixes from Spark UI symptoms, cluster settings, and query patterns.
by alirezarezvani
chief-data-officer-advisor is a strategic CDO skill for startup data decisions: AI training data rights, warehouse vs lakehouse vs mesh strategy, customer-data asset valuation, M&A readiness, and data team hiring. Includes references and Python tools for decision support, not tactical data engineering.
by alirezarezvani
cdo-review is a Chief Data Officer review skill for pressure-testing data strategy plans, AI training data rights, architecture choices, data productization, M&A diligence, and data-team hiring before commitments.
by markdown-viewer
The data-analytics skill creates PlantUML diagrams for data analysis workflows, including ETL, ELT, data lakes, warehouses, streaming pipelines, log analytics, and BI dashboards. It is optimized for clear source-to-destination flow, AWS analytics/database stencils, and practical data-analytics guide output—not generic software or cloud architecture diagrams.
by tinybirdco
tinybird-python-sdk-guidelines helps you install and use tinybird-sdk for Python-based Tinybird projects. It covers datasources, endpoints, clients, connections, migration from legacy files, and backend development workflows with build and deploy guidance.
by K-Dense-AI
The lamindb skill helps you work with LaminDB, an open-source biology data framework for making data queryable, traceable, reproducible, and FAIR. Use it for lamindb for Data Analysis, metadata curation, ontology-based annotation, schema validation, and lineage-aware workflows across notebooks and pipelines.