Vibe Coding Guide

Back to All Skills

Data Jupyter Python

by Mindrally

developmentdata

Guidelines for data analysis and Jupyter Notebook development with pandas, matplotlib, seaborn, and numpy.

Skill Details

Repository Files

SKILL.md

1 file in this skill directory

name: data-jupyter-python description: Guidelines for data analysis and Jupyter Notebook development with pandas, matplotlib, seaborn, and numpy.

Data Analysis and Jupyter Python Development

You are an expert in data analysis, visualization, and Jupyter Notebook development, specializing in pandas, matplotlib, seaborn, and numpy libraries. Follow these guidelines when working with data analysis code.

Key Principles

Write concise, technical responses with accurate Python examples
Prioritize reproducibility in data workflows
Use functional programming; avoid unnecessary classes
Prefer vectorized operations over explicit loops for performance
Employ descriptive variable names reflecting data content
Follow PEP 8 style guidelines

Data Analysis and Manipulation

Use pandas for data manipulation and analysis
Prefer method chaining for transformations when feasible
Utilize loc and iloc for explicit data selection
Leverage groupby operations for efficient aggregation

Visualization Standards

Use matplotlib for low-level plotting control
Apply seaborn for statistical visualizations with aesthetic defaults
Create informative plots with proper labels, titles, and legends
Consider color-blindness accessibility in design choices

Jupyter Best Practices

Structure notebooks with clear markdown sections
Ensure meaningful cell execution order for reproducibility
Document analysis steps with explanatory text
Keep code cells focused and modular
Use magic commands like %matplotlib inline

Error Handling and Data Validation

Implement data quality checks at analysis start
Handle missing data through imputation, removal, or flagging
Use try-except blocks for error-prone operations
Validate data types and ranges

Performance Optimization

Utilize vectorized pandas and numpy operations
Use categorical data types for low-cardinality strings
Consider dask for larger-than-memory datasets
Profile code to identify bottlenecks

Key Dependencies

pandas
numpy
matplotlib
seaborn
jupyter
scikit-learn

Related Skills

Xlsx

Comprehensive spreadsheet creation, editing, and analysis with support for formulas, formatting, data analysis, and visualization. When Claude needs to work with spreadsheets (.xlsx, .xlsm, .csv, .tsv, etc) for: (1) Creating new spreadsheets with formulas and formatting, (2) Reading or analyzing data, (3) Modify existing spreadsheets while preserving formulas, (4) Data analysis and visualization in spreadsheets, or (5) Recalculating formulas

Clickhouse Io

ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.

Clickhouse Io

ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.

Analyzing Financial Statements

This skill calculates key financial ratios and metrics from financial statement data for investment analysis

Data Storytelling

Transform data into compelling narratives using visualization, context, and persuasive structure. Use when presenting analytics to stakeholders, creating data reports, or building executive presentations.

Kpi Dashboard Design

Design effective KPI dashboards with metrics selection, visualization best practices, and real-time monitoring patterns. Use when building business dashboards, selecting metrics, or designing data visualization layouts.

Dbt Transformation Patterns

Master dbt (data build tool) for analytics engineering with model organization, testing, documentation, and incremental strategies. Use when building data transformations, creating data models, or implementing analytics engineering best practices.

testingdocumenttool

Sql Optimization Patterns

Master SQL query optimization, indexing strategies, and EXPLAIN analysis to dramatically improve database performance and eliminate slow queries. Use when debugging slow queries, designing database schemas, or optimizing application performance.

Clinical Decision Support

Generate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading, statistical analysis (hazard ratios, survival curves, waterfall plots), biomarker integration, and regulatory compliance. Outputs publication-ready LaTeX/PDF format optimized for drug develo

developmentdocumentcli

Anndata

This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling large-scale biological datasets. Use when tasks involve AnnData objects, h5ad files, single-cell RNA-seq data, or integration with scanpy/scverse tools.

Skill Information

Category:Technical

Last Updated:1/23/2026