Hi, I'm

Dan Doughty.

Director of DevOps / Infrastructure / Platform Engineering

I direct DevOps and SRE engineering across multi-cloud healthcare infrastructure — 10 hosted products spanning two business units, under HIPAA, SOC, and FDA compliance. I build teams, not just systems: growing and leading through team leads, formalizing roadmap and delivery process, and driving the automation and observability work that cuts incident load without adding headcount.

Dan Doughty — Director of DevOps / Infrastructure / Platform Engineering profile image

About

I’m a Senior Staff SRE at WellSky, directing DevOps and SRE work across two engineering teams (through two team leads — one technical, one managerial) supporting 10 hosted healthcare products across two business units. That includes the FDA-regulated WellSky Transfusion appliance, running under HIPAA and SOC compliance. I manage $5M in annual cloud spend across Azure and GCP, and led disaster recovery across 8 product lines during the CrowdStrike outage, restoring all production services within 18 hours.

Before WellSky, I founded and commanded an electronic-warfare counter-drone unit supporting Ukrainian defense operations (UA Owls), led SRE teams at Direct Supply, and spent nearly 8 years across three roles at Epiq Systems — including a stretch as Acting Director of Compute Engineering co-leading response to a multi-datacenter ransomware intrusion.

I’m targeting Director-level DevOps / Infrastructure / Platform Engineering roles, 100% remote.

Stack:
  • Terraform (Spacelift-managed)
  • Kubernetes (AKS, Rancher)
  • Azure
  • AWS
  • GCP
  • Python
  • Perl
  • Bash / PowerShell
  • New Relic
  • Grafana
  • Splunk
  • PagerDuty
  • CI/CD

Experience

Senior Staff Site Reliability Engineer - WellSky
Mar 2023 - present

Lead two DevOps/SRE teams through two team leads (one technical, one managerial), one per business unit, supporting Linux and Windows application stacks across multi-region Azure, a Phoenix colocation datacenter (VMware), and GCP — 10 hosted healthcare products across two business units, under HIPAA and SOC compliance (plus FDA for the Transfusion appliance).

  • Peaked at 13 engineers across 3 team leads and 3 business units for roughly a year.
  • Formalized the team’s roadmap process for reliable stakeholder-driven delivery.
  • Delivered WellSky Transfusion’s first cloud-native release on a tight deadline — an FDA-regulated appliance, first time the product ran in the cloud.
  • $5M in annual cloud spend across Azure and GCP, managed via negotiated per-environment service tiers and an ongoing utilization review cadence.
  • Directed DR testing across 8 product lines; restored all production services within 18 hours during the CrowdStrike outage.
  • Reduced database-incident blast radius from 17 clients per incident to 1-2 through tenant isolation work.
  • Cut Code Red incidents and MTTR by ~50%; improved application responsiveness by 60% in the Community Services product via targeted performance tuning.
  • Onboarded and fully trained 4 new team members; standardized project reporting and repeatable DevOps practices for provisioning, patching, and monitoring.
  • Modernized Rancher-managed Kubernetes (dev/prod) and AKS clusters — health checks, isolated node pools, self-healing deployments; active blue/green Kubernetes upgrades in Phoenix and Azure.
Founder / Commander - UA Owls
Apr 2022 - Feb 2023

Founded and commanded an electronic-warfare counter-drone unit supporting Ukrainian defense operations.

  • Synthesized input from product experts and theater experts into a strategic call: most allied counter-drone units were fixated on RF jamming — a downstream tactic, easily adapted around — rather than disrupting the kill chain further upstream through direct targeting of enemy drone operators and disruption of production capacity. Advocated for and trained toward the upstream approach; this is roughly the direction counter-drone doctrine across the war ultimately took.
  • Assessed Ukraine’s highest-priority counter-drone capability gaps directly from unit-level battlefield conditions, translating that into concrete equipment and training requirements.
  • Evaluated the available electronic-warfare/counter-drone product landscape against those requirements.
  • Evaluated strategic/theater-level organizations versus regimental-level units as delivery targets, and chose regimental delivery — the level positioned to act fastest on the data the systems produced.
  • Raised funds and ran procurement to source assessed equipment; managed delivery into theater.
  • Operated in a deliberately ambiguous environment: the U.S. State Department tolerated but didn’t endorse the effort, and periodically steered toward lower-risk courses of action. Made independent, calculated decisions to proceed where battlefield judgment diverged from that guidance.
  • Personally delivered training to 3 Ukrainian units over 9 months — the scale sustainable without further fundraising — on equipment operation and the upstream kill-chain framing above, and cross-trained a team member to extend delivery capacity beyond himself.
Site Reliability Engineering Manager - Direct Supply
Jul 2021 - Apr 2022

Led a 6-person SRE team supporting Linux and Windows production workloads across hybrid AWS/GCP.

  • Built Terraform-driven New Relic observability, modular and versioned, adoptable by product dev teams for self-service reliability.
  • Formalized a liaison program between SRE and product development teams — improved observability adoption and accelerated feature deployments.
  • Groomed an SLO/SLI backlog covering the four golden signals.
Site Reliability Engineering Manager - Epiq Systems
Jun 2018 - Jun 2021

Managed a 6-person SRE team delivering application-centered, full-lifecycle Linux solutions.

  • Operationalized Epiq’s first cloud-native application (Epiq Discovery) — auto-scaling Ubuntu farm, Aurora Postgres, Kafka, ALB/ELB, S3, VPCs. Engineered an A/B DB pattern that cut CD pipeline deploy time by 90%.
  • Ran 300TB Oracle RAC clusters plus MySQL, Aurora, Postgres, and ElasticSearch.
  • Spent the first half of 2020 as Acting Director of Compute Engineering, co-leading response to a multi-domain, multi-datacenter ransomware intrusion — coordinated detection, containment, eradication, and recovery across AD and Server Engineering.
  • Datacenter convergence eliminated 75% of leased rack space across 4 new colos.
  • Partnered with sales on the Brainspace Discovery 5 offering, contributing to roughly $50M in revenue; replaced a 3-day VMware build with a 15-build/day Vagrant + GitLab recipe.
Supervisor, Linux Engineering
Mar 2014 - Jun 2018

Managed an international team of four IT professionals supporting Linux application and OS work, troubleshooting, patching, and code deployments.

  • Centrify cost-justification analysis for legacy systems drove $250k in annual OPEX savings.
  • Designed a Change Advisory Board process that unblocked non-production updates and minimized development downtime.
  • Visualized KPIs via open-source tooling, turning capacity-planning answers from weeks into hours.
Senior Linux Engineer
Sep 2013 - Mar 2014
Led a monitoring vision where KPIs notify on-call via a tiered notification system, at zero incremental OPEX using open-source tooling.
Apr 2011 - Sep 2013

Ran all facets of a Kansas City IT business: strategic planning, operations, merchandising, marketing, customer relations, and financial management.

  • Built a tablet-based GPS navigation solution for snow-removal truck drivers, replacing hard-copy binders — paid for itself in its first season.
Systems Administrator - NIC
Apr 2008 - Apr 2011
Delivered Linux enterprise solutions across two remote sites on open-source technology; Oracle DB and Solaris Zones for ACH and credit-card processing under strict PCI/DSS adherence.
Senior Systems Administrator - Sprint
Apr 2001 - Apr 2007

Supported FCC-reportable, geographically redundant Messaging and Home Location Registry services across 20+ locations.

  • Negotiated vendor contract requirements enabling millisecond text-message billing — a 4% increase in billable messages, worth $72M in the first year, at no new software cost.

Education

2000 - 2010
BA, Slavic Languages & Literature
University of Kansas
1995 - 1997
Russian
Defense Language Institute Foreign Language Center

Outcomes

99.99% uptime
Sustained across WellSky's multi-cloud healthcare infrastructure.
~50% MTTR / Code Red reduction
Via automation, observability, and network tuning at WellSky.
18-hour full restore
Directed disaster recovery across 8 product lines during the CrowdStrike outage.
17 clients to 1-2 per incident
Reduced database-incident blast radius through tenant isolation work.
90% faster CD pipeline
Engineered an A/B database pattern for Epiq Discovery's deploy pipeline.
$50M revenue contribution
Partnered with sales on the Brainspace Discovery 5 rollout at Epiq.

Contact

Selectively exploring Director-level DevOps / Infrastructure / Platform Engineering roles, 100% remote.