Case StudyHealthcare Provider

Hospital Readmission Prediction

A machine learning pipeline that predicts hospital readmissions within 30 days using patient encounter data, featuring advanced feature engineering on ICD-10 codes and ensemble model comparison.

ML
Healthcare
Hospital Readmission Prediction

Project Overview

This project predicts hospital readmissions within 30 days using patient encounter data. It includes end-to-end preprocessing, feature engineering, exploratory data analysis, model training, and prediction generation.

The Problem

Hospital readmissions within 30 days are a major cost driver and quality indicator in healthcare. Identifying which patients are at risk of returning before they leave allows hospitals to intervene with targeted discharge planning, follow-up scheduling, and care coordination.

The Goal

Build a robust prediction pipeline that handles high-dimensional clinical data, engineers meaningful features from raw ICD-10 codes and procedure records, and compares multiple classification models to find the best performer on an imbalanced dataset.

Dataset

Each row represents one hospital encounter. Key features include:

  • Admission and discharge dates for computing stay duration
  • 25 diagnosis code columns (ICD-10) and 25 procedure code columns per encounter
  • Admission type, source, and discharge status codes for clinical context
  • DRG codes for grouping encounters into clinical categories
  • Target variable: Readmitted within 30 days (binary)

The dataset presented three core challenges: high dimensionality from dozens of sparse code columns, significant missing values across most features, and severe class imbalance in the target variable.

Methodology

1. Preprocessing

  • Removed patient ID, handled NaN values, replaced placeholder dashes with zeros
  • Computed stay duration from admission and discharge dates
  • Added a binary Long Stay flag for encounters exceeding the threshold

2. Feature Engineering

  • Counted total diagnoses and procedures per encounter
  • One-hot encoded admission type, source, and discharge status codes
  • Mapped DRG and diagnosis codes to higher-level clinical categories
  • Created binary flags for the top 200 most frequent diagnosis codes
  • Converted all boolean features to integers for model compatibility

3. Exploratory Data Analysis

  • Visualized relationships between stay length, diagnosis counts, and readmission rates
  • Examined class distribution and feature correlation patterns

4. Modeling

  • Train/test split at 80/20
  • Trained and compared: XGBoost, LightGBM, Random Forest, and multiple Naive Bayes variants

5. Evaluation

  • Metrics: Accuracy, ROC AUC, Precision, Recall, and F1 Score
  • Most models achieved high accuracy but struggled with recall for the minority class, a common pattern with imbalanced clinical data

Results

High overall accuracy across all models, but low recall for the readmission class due to class imbalance. Random Forest was selected for the final submission based on its balance of precision and generalization. Discrepancies between models (particularly Random Forest vs. LightGBM) highlight the inherent difficulty of predicting readmission from encounter-level features alone.

Key Learnings

  • Class imbalance dominates model behavior. Without resampling or cost-sensitive learning, even strong models default to predicting the majority class
  • Feature engineering matters more than model selection. Mapping raw ICD-10 codes to clinical categories and counting diagnoses per encounter added more predictive power than switching between algorithms
  • Ensemble models are not always the answer. On heavily imbalanced data, a simpler model with well-tuned thresholds can outperform a complex ensemble that optimizes for overall accuracy

Tech Stack

  • Language: Python 3.10+
  • Libraries: pandas, NumPy, scikit-learn, XGBoost, LightGBM, matplotlib, seaborn
  • Environment: Jupyter Notebook / Google Colab
building something exciting?

We want to hear from you

Join 50+ businesses scaling with AI.

Powered By

Supabase
Vercel
Next.js
OpenAI
Anthropic
Stripe
Twilio
HubSpot
Zapier
Make
Supabase
Vercel
Next.js
OpenAI
Anthropic
Stripe
Twilio
HubSpot
Zapier
Make
Supabase
Vercel
Next.js
OpenAI
Anthropic
Stripe
Twilio
HubSpot
Zapier
Make
Supabase
Vercel
Next.js
OpenAI
Anthropic
Stripe
Twilio
HubSpot
Zapier
Make

Book a technical discovery

Tell us about your operational bottlenecks. We'll engineer an autonomous system to eliminate them entirely.