Zero Tolerance for Failure: Rebuilding a Seven-Hospital ETL Pipeline Mid-Crisis
How APSIS stopped nightly Caboodle ETL failures that were cascading into reporting outages across an entire hospital network.
The situation
The data team at a seven-hospital network had been fighting the same war every morning for eleven weeks. The Caboodle ETL — the overnight pipeline that loads Epic clinical and operational data into the enterprise analytics environment — was failing. Not occasionally. Routinely.
When the ETL failed, it did not fail quietly. Downstream reporting jobs that depended on fresh data simply did not run. Quality dashboards serving seven hospitals went dark. Revenue cycle reports that should have been in leadership inboxes by 7:00 AM were not there. Clinical operations teams were making decisions on day-old data. In a seven-hospital system where executives rely on near-real-time analytics for capacity management and throughput decisions, day-old data is not a minor inconvenience. It is an operational problem.
The organization’s internal team had applied three rounds of fixes. None held. They had opened a high-priority case with Epic; the support ticket was active but moving slowly. They needed someone who could walk in, diagnose the root cause rather than the symptoms, and rebuild the pipeline to a standard it had not had before.
We’d been fighting these failures for nearly three months. Our team was exhausted. The Epic ticket was going nowhere fast. What APSIS brought was a completely different diagnostic approach — they weren’t looking at what failed, they were asking why it was designed the way it was.
Challenge and approach
The challenge
- 11 consecutive weeks of nightly ETL failures across a 7-hospital pipeline
- TempDB running as a single data file, creating severe I/O contention during the high-volume ETL window
- Parallelized job design with implicit dependencies — downstream jobs starting before their input data was ready, writing corrupt or incomplete records
- Silent data integrity gaps: partial-load runs completing without error codes but writing incomplete data sets
- No alerting infrastructure — failures discovered when clinical operations staff called to report missing reports
The APSIS approach
- Deployed an APSIS Caboodle/SQL Server specialist with Epic ETL architecture experience
- Rebuilt TempDB across eight equal data files per SQL Server best practice, eliminating the I/O bottleneck
- Redesigned ETL job sequencing with explicit dependency chaining — no parallel execution without confirmed prerequisites
- Conducted a full data integrity audit across Caboodle; identified and corrected 23 tables with inconsistent records from partial loads
- Implemented SQL Server Agent alerting and a custom monitoring dashboard surfacing job status, runtime trends and failure prediction
The broader clinical impact
Data infrastructure failures are rarely just IT problems. When the Caboodle ETL failed, the downstream effects touched clinical operations directly: capacity managers were flying blind, throughput decisions were based on yesterday’s census, and quality reporting gaps created risk for an organization under continuous regulatory scrutiny.
The stabilization APSIS delivered was not just a technical fix. It restored the analytical foundation the health system’s clinical and operational leadership depends on to run seven hospitals. That is a different kind of outcome — and it is the kind that only comes when the team solving the problem understands both the technical architecture and what the data is actually used for.
Results
The night the rebuilt pipeline ran for the first time, every job completed successfully. The following morning, quality dashboards were populated, revenue cycle reports were in leadership inboxes before 7:00 AM, and the data team — which had been on morning crisis calls for eleven weeks straight — had a normal day.
The monitoring infrastructure APSIS implemented changed the operational posture entirely. Instead of discovering failures after clinical staff called in, the data team now receives alerts within 60 seconds of a job anomaly. Trend data on runtime allows them to anticipate degradation before it becomes failure.
The seven-hospital network retained APSIS for ongoing Caboodle managed services support following the engagement — a direct result of the confidence built during the stabilization work.
APSIS capabilities demonstrated
Cloud & database infrastructure
- Epic Caboodle architecture and ETL design
- SQL Server TempDB engineering and I/O optimization
- ETL job sequencing and dependency architecture
- Data integrity audit and remediation
- Monitoring framework design and implementation
Managed services capability
- Ongoing Caboodle managed services — retained post-engagement
- 24-hour monitoring with proactive alerting
- Root cause diagnosis vs. symptom treatment
- No disruption to active clinical operations during work
- Retained as long-term data infrastructure partner
This page publishes the full engagement record. The downloadable PDF is the same case study in the client-facing format, with identical figures.
Facing something similar?
APSIS responds to inbound inquiries within four business hours, routes a request for a senior VP-level advisor within two business hours, and deploys credentialed healthcare IT professionals on a 48-hour SLA.
Other case studies
- CS 01From Two Hours to Five Minutes: Restoring Epic Reporting for a Regional Health System96% reduction in report runtime
- CS 0314 Analysts, 48 Hours, One Chance to Get It Right: Epic Go-Live Surge Support at an Academic Medical Center62% reduction in post-live help desk tickets
- CS 04The 24-Hour Bench: How APSIS Won a Competitive Staffing Race for a Prime Contractor100% placement, 100% converted to full-time