Comprehensive leadership experience in architecture, design, and development of large-scale data center and edge infrastructure for highly performant, resilient, and distributed applications within Telco/Service Provider, Video Streaming, and FinTech/Blockchain industries.
- Comprehensive leadership experience in architecture, design, automation, and AI-enabled operations of large-scale distributed infrastructure for highly performant, resilient applications within Telco/Service Provider, Video Streaming, and FinTech/Blockchain industries.
- Accomplished technical leader known for building high-performing teams, optimizing workflows, and fostering innovation. Demonstrated expertise in team management & motivation, talent acquisition, and capital budgeting in alignment with organizational goals.
- Proven leadership track record of building strong internal/external customer relationships, driving cross-functional collaboration, forging strategic vendor partnerships, and delivering exceptional results.
- Proven ability to leverage hands-on cross-domain expertise to quickly comprehend complex landscapes and develop innovative solutions tailored to the unique challenges and requirements of new domains and industries.
Production engineering and infrastructure leader with 25+ years building, standardizing, securing, and operating high-stakes distributed systems across fintech, blockchain, telecom, video streaming, hybrid-cloud, and bare-metal environments.
- Deep hands-on experience with Terraform, Ansible, GitHub Actions, GitLab CI, Linux, AWS, Kubernetes, PostgreSQL, HashiCorp Vault, Datadog, Grafana, Splunk, network architecture, and production incident response.
- Known for incremental modernization: replacing fragile manual operations with automation, observability, security defaults, runbooks, and change patterns that keep production systems running.
- Built production environments from scratch, standardized deployment workflows, reduced MTTR from hours to minutes, and improved deployment speed without disrupting business-critical systems.
- Led teams and programs across rapid-growth startups and large enterprises, translating architecture decisions across engineering, security, operations, vendors, and business stakeholders.
Senior SRE and platform infrastructure leader with 25+ years designing, automating, observing, and operating reliable Linux, Kubernetes, bare-metal, cloud, and data-center systems across blockchain, fintech, telecom, and service-provider environments.
- Deep hands-on experience with Linux/Unix production systems, Kubernetes, Terraform, Ansible, GitHub Actions, GitLab CI, Python/Bash automation, KVM, HashiCorp Vault, Datadog, Prometheus/Grafana, Splunk, Dynatrace, and incident response.
- Designed and built automation, runbooks, telemetry, dashboards, and deployment workflows that replaced fragile manual operations and made large fleets more reliable, repeatable, and observable.
- Designed, implemented, and automated bare-metal, COLO, and cloud infrastructure at scale: 1,000+ server fleets, 60+ blockchain networks, 13 cloud/COLO providers, multi-site container platforms, and production data-center migrations.
- Comfortable working across reliability engineering, platform architecture, data-center hardware, networking, storage, security controls, and application teams to turn operational pain into durable systems.
Principal production engineering and platform leader with 25+ years designing, building, hardening, and operating reliable workflow platforms across media/video streaming, telecom, fintech/blockchain, and enterprise environments, with current hands-on AI/LLM, MCP, RAG, and observability work.
- Built production-grade solution layers with Python and automation, API/service integration, Kubernetes, Terraform/Ansible, CI/CD, telemetry, dashboards, runbooks, rollback patterns, and incident/postmortem remediation.
- Owned design & implementation of reference architectures and production patterns for high-impact platforms: Verizon OnCue OTT/IPTV across 1000+ edge locations and 15+ Pbps, Verizon Connect TCCP at 100% uptime over 1.5+ years, Cube.Exchange's high-security production environment from scratch, and Nirvana agent/MCP integrations that cut MTTR from 4+ hours to <15 minutes.
- Translated ambiguous stakeholder needs into maintainable systems by defining boundaries, data flows, deployment patterns, operational controls, and reusable scaffolding across cloud, COLO, bare-metal, secure enterprise, and constrained network environments.
- Hands-on AI workflow experience includes MCP integrations, LangGraph/FastMCP retrieval layers, LiteLLM gateway patterns, vector RAG bake-offs, OpenTelemetry tracing, Prometheus/Grafana observability, Langfuse evals, and budget/quality controls.
Hands-on Staff-level SRE and platform engineer with deep experience designing, automating, observing, and operating highly available cloud-native systems on Amazon EKS, Kubernetes, and Linux at scale. 25+ years across fintech, blockchain, telecom, and high-throughput data environments - spanning cloud, COLO, and high-density bare-metal - with current Kubernetes, observability, and LLM-infrastructure work.
- Deep hands-on SRE/infra execution: Amazon EKS and Kubernetes clusters, Docker, Terraform/Ansible infrastructure-as-code, GitHub Actions and GitLab CI/CD, Prometheus/Grafana, Datadog, OpenTelemetry, Splunk, and production incident response across AWS and multi-provider fleets.
- Designed, built, and operated highly available, scalable, secure cloud-native infrastructure: 1,000+ server fleets, multi-site Active-Active container platforms with 100% uptime over 1.5+ years, and zero-touch provisioning that cut deploy cycles from 30 days to 8 hours.
- Built observability from the ground up - telemetry ingestion, dashboards, alerting, and tracing (Prometheus/Grafana, Datadog/Vector, OpenTelemetry, Splunk/Dynatrace) - that replaced fragile manual operations and reduced MTTR from 4+ hours to under 15 minutes.
- Built and operated the infrastructure behind high-throughput telemetry systems: 1.5M TPS and 120TB/day ingest with hot/cold tiering (600TB Postgres/Citus, 11PB Dremio over S3) and Redshift–Iceberg/Parquet federation.
- Built infrastructure and CI/CD for an internal LLM platform: LiteLLM OpenAI-compatible gateway, OTel tracing, virtual keys, caching, guardrails, and budget controls - plus EKS/Kubernetes autoscaling runner platforms for pipeline parallelism and isolation.
- Drives reliability as a force multiplier - runbooks, IaC standards, and code review that make IC work higher-leverage - while collaborating across security, networking, storage, and application teams.
Infrastructure engineering manager with 25+ years building and leading teams across cloud, Kubernetes/container platforms, networking, storage, security, edge, and compute infrastructure in telecom, fintech/blockchain, service-provider, and startup environments.
- Led infrastructure teams building and operating production platform systems at scale: 1,000+ server fleets across 13+ regions/providers, 60+ blockchain networks on 1,300+ servers, Verizon router telemetry designed for 1.5M TPS and 120TB/day, and nationwide video platforms across 1,000+ edge locations and 15+ Pbps.
- Designed and built platform foundations: Active-Active container cloud with eBGP Clos networking, Docker Swarm, Consul DNS, Ceph storage, 100% uptime over 1.5+ years; Kubernetes clusters with autoscaling GitHub Actions runners; and k3s/Cilium ingress patterns for secure service exposure and observability.
- Built self-service, reproducible infrastructure with Terraform, Ansible, GitHub/GitLab CI, zero-touch provisioning, reusable roles, runbooks, HashiCorp Vault, PostgreSQL automation, and cost/capacity analysis across AWS, COLO, bare-metal, and hybrid-cloud environments.
- Hired, coached, and grew infrastructure teams through rapid scaling: 10 hires and 5 promotions at Figment, 8+ DevOps engineers and contractors at Verizon Connect, and 3 senior engineers at Nirvana Labs.
Engineering manager and platform leader with 25+ years building high-throughput, low-latency, cost-sensitive distributed systems, plus current hands-on LLM gateway, agent, RAG, observability, and evaluation infrastructure.
- Built inference-adjacent LLM platform components: LiteLLM as an OpenAI-compatible gateway across local and hosted providers with Langfuse callbacks, OTel export, virtual keys, caching, guardrails/policies, and budget controls.
- Designed high-throughput distributed platforms and data paths: Verizon Product Analytics Cloud for 20M routers, 200M Wi-Fi clients, 1.5M TPS, and 120TB/day ingest; OnCue OTT/IPTV across 1,000+ edge locations and 15+ Pbps; and Cube.Exchange's high-security, low-latency production environment from scratch.
- Evaluated high-density accelerator/GPU and video-encoding platforms in a lab for objective quality/performance analysis, including Dell C410x testing with 16x NVIDIA GPUs; found the full chassis overloaded PCIe, split it across two servers at 8 GPUs/server, and drove $9M+ in encoding cost savings plus $200K+ in vendor-funded lab equipment savings.
- Balanced cost, performance, reliability, and capacity in gray-area infrastructure decisions: $80M hardware/vendor savings, Redshift-to-Dremio/OpenShift architecture planning, spare-capacity analytics, blockchain API performance remediation, and multi-provider/region rollouts.
- Led and grew strong infrastructure/platform teams through hands-on technical direction, code review, hiring, coaching, and high-leverage automation across startups and large enterprises.
Engineering leader with 25+ years building and scaling globally distributed teams that ship reliable software on DevSecOps, CI/CD, and platform systems across telecom, fintech/blockchain, and startup environments, now driving AI-native developer-productivity practices into engineering workflows, code review, and continuous delivery.
- Led globally distributed engineering and DevOps organizations across telecom, fintech/blockchain, and startups: a 5-person node DevOps team running 50+ L1 and 75+ L2 blockchains on 1,000+ servers across 13+ regions at Nirvana Labs, and a software/data engineering org building Verizon's Product Analytics Cloud (20M routers, 1.5M TPS, 120TB/day).
- Drove operational excellence as a contract, not an aspiration: TCCP Active-Active container platform held 100% uptime over 1.5+ years and cut deploy cycles from 30 days to 8 hours; standardized runbooks, telemetry, and incident practices at Nirvana cut MTTR from 4+ hours to <15 minutes across a 125+ chain fleet on 1,000+ servers and 13+ regions, and led automation across 60+ networks and 1,300+ servers at Figment.
- Championed AI-native engineering and developer productivity: established reusable agent standards, MCP integrations, and AI-assisted PR reviews driving 5x faster role development; built a FastMCP retrieval layer, a LangGraph docs RAG agent, an LLM text-to-SQL analytics assistant, and a LiteLLM gateway with guardrails, caching, and budget controls.
- Modernized delivery onto modular, independently deployable platforms and sequenced production migrations: decomposed deployments onto container/Kubernetes platforms (Docker Swarm, k3s/Cilium, eBGP Clos networking, Ceph) with end-to-end GitLab and GitHub Actions CI/CD, zero-touch provisioning, and AWS-managed-service to self-hosted OpenShift migration planning.
- Hired, coached, and grew engineering teams through rapid scaling (10 hires / 5 promotions at Figment, 8+ engineers building the first Telematics DevOps team at Verizon Connect, 3 senior engineers at Nirvana Labs), and led through distinguished engineers, contractor teams, and vendor/SOW relationships while partnering with executive, product, finance, and security stakeholders in remote-first, documentation-heavy cultures.
Hands-on platform and DevOps engineer with 25+ years building CI/CD, infrastructure automation, reproducible environments, internal tooling, observability, and developer workflows that help teams build, test, and ship faster.
- Built CI/CD and feedback-loop systems: GitHub Actions PR preview environments validating sync, OS/chain metrics, and dashboards; GitHub Actions Runner Controller autoscaling runners; and Rust application pipelines with unit tests and static analysis.
- Built reproducible environments and platform scaffolding with macOS/Colima/Lima, Chezmoi, Homebrew Bundle, mise, pre-commit validation, Lima test VMs, scripted k0s/k3s/kubeadm clusters, Terraform modules, Ansible roles, and reusable runbooks.
- Improved developer and operator workflows with reusable agents, MCP integrations, observability/agent tracing, AI-assisted PR reviews, internal Grafana dashboards, documentation, onboarding, and code review; drove 5x faster new-network role development and MTTR from 4+ hours to <15 minutes.
- Designed workflow automation that tightened feedback loops: Verizon Connect TCCP reduced deploy cycles from 30 days to 8 hours, Redshift performance scripts established query benchmarks, and PDT replaced manual Sneakernet patching with zero-touch deployments in a highly secure Windows server environment, reclaiming 500+ engineer-hours annually.
Infrastructure and platform engineering director with 15+ years designing, building, and leading high-throughput distributed systems - from a nationwide 15+ Pbps video platform to multi-cloud blockchain infrastructure spanning 125+ chains, 1,000+ servers, and 13+ providers. Deep technical fluency in cloud infrastructure, Kubernetes, and distributed systems at scale, paired with direct experience owning mission-critical blockchain and Web3 infrastructure as an engineering leader at Figment and Nirvana Labs.
- Director-level engineering leader with blockchain and Web3 infrastructure depth: led a 5-person global node DevOps team at Nirvana Labs operating 50+ L1s, 75+ L2s, 125+ chains, and 1,000+ servers across 13+ regions and 6 cloud/bare-metal providers; previously Engineering Manager at Figment.io automating 60+ blockchain networks across 1,300+ servers on 13 cloud/COLO providers.
- Designed and operated high-throughput, mission-critical distributed systems at scale: Verizon OnCue OTT/IPTV nationwide platform (1,000+ edge locations, 40+ POPs, 15+ Pbps); Product Analytics Cloud router telemetry (20M routers, 200M Wi-Fi clients, 1.5M TPS, 120TB/day); and a multi-site Active-Active container cloud platform sustaining 100% uptime over 1.5+ years.
- Built and owned multi-cloud infrastructure strategy and cost efficiency: standardized Ansible/Terraform platform across AWS, DigitalOcean, Latitude, OVH, Servers.com, and Nirvana Cloud; drove hardware evaluation and architecture decisions yielding $80M+ in capital cost savings; led AWS-managed-services-to-self-hosted OpenShift migration to reduce opex and vendor lock-in.
- Hired, developed, and retained engineering leaders and ICs through rapid scaling: 10 hires and 5 promotions at Figment, 3 senior engineers at Nirvana Labs; defined hiring strategy, owned milestone tracking, and established documentation, code review, and onboarding standards as both engineering manager and head of nodes.
- Set platform technical direction and drove execution excellence with measurable outcomes: reduced new-network provisioning from 15+ days to under 24 hours, fleet-wide upgrades from multi-day efforts to under 1 hour, and MTTR from 4+ hours to under 15 minutes - through structured Ansible/Terraform roles, runbooks, and CI/CD with PR preview validation.
- Hands-on technical depth in Kubernetes (k3s, Cilium, autoscaling GitHub Actions Runner Controller), AWS, Terraform/Ansible IaC, eBGP Clos networking, Ceph/ZFS storage, HashiCorp Vault, Datadog/Prometheus/Grafana observability, and zero-touch provisioning across bare metal, COLO, and cloud.
Principal-level infrastructure engineer with 25+ years of hands-on Linux systems engineering, specializing in zero-touch bare-metal provisioning and full infrastructure lifecycle management across large-scale fleets - 1,000+ servers across six providers at Nirvana Labs and 1,300+ across 13 cloud/COLO providers at Figment. Builds Prometheus/Grafana fleet observability and automated, self-healing operations (Terraform/Ansible, Python/Go) that drive MTTR from hours to minutes. Brings hardware-layer depth: direct NVIDIA GPU and PCIe-topology characterization (GPU lab) plus high-density OCP-rack, eBGP Clos, and high-throughput, high-performance infrastructure.
- Designed and automated zero-touch bare-metal provisioning and full infrastructure lifecycle management at scale - 1,000+ servers across six cloud and bare-metal providers at Nirvana Labs and 1,300+ across 13 cloud/COLO providers at Figment - with Terraform/Ansible, structured inventories, reusable roles, and runbooks, cutting new-environment provisioning from 15+ days to under 24 hours.
- Automated Linux infrastructure with Go and Python: an open-source Go Terraform provider driving bare-metal Dell PowerEdge configuration and deployment through Redfish/BMC REST APIs - infrastructure-as-code reaching the server hardware and firmware-management layer - plus a Python/Ansible framework for PostgreSQL and fleet operations.
- Built fleet-health observability from the ground up with Prometheus and Grafana - telemetry ingestion, dashboards, and alerting giving SRE, support, and engineering shared visibility into node and fleet operations - alongside large-scale Datadog, Splunk, and Dynatrace deployments for metrics, logs, and application performance monitoring.
- Cut MTTR from 4+ hours to under 15 minutes through automated self-healing operations: remediation runbooks, ZFS snapshot/restore with NVMe 4K block-size alignment for fast node recovery and high-throughput disk I/O, and CI/CD with PR preview environments validating sync, OS, and host metrics before production.
- Engineered high-density compute and fabric foundations: a multi-site Active-Active container platform on high-density OCP racks with eBGP Clos networking, Docker Swarm, and Ceph storage sustaining 100% uptime over 1.5+ years.
- Brings direct NVIDIA GPU and PCIe-topology hardware depth - built a GPU video-encoding lab characterizing a Dell C410x with 16x NVIDIA GPUs, diagnosed a PCIe over-subscription bottleneck, and re-architected to 8 GPUs/server, supporting $9M+ in encoding cost savings.
Senior cloud infrastructure engineer with 25+ years of hands-on Linux systems engineering - designing and operating multi-region, highly available infrastructure and automating deployments into partner- and customer-managed environments. Ran fleets of 1,000+ servers across six cloud and bare-metal providers in 13+ regions (Nirvana Labs) and 1,300+ servers across 13 cloud/COLO providers (Figment); designed a multi-site Active-Active platform with automated failover sustaining 100% uptime over 1.5+ years (Verizon). Runs Terraform/Ansible infrastructure-as-code and git-driven operations end to end, builds observability from the ground up (Prometheus/Grafana, Datadog, OpenTelemetry, Splunk), and drives MTTR from 4+ hours to under 15 minutes.
- Designed and operated multi-cloud, multi-region, highly available infrastructure at fleet scale: 125+ blockchain networks on 1,000+ servers across six cloud and bare-metal providers (AWS, DigitalOcean, Latitude, OVH, Servers.com, Nirvana Cloud) and 13+ regions, standardized on one Terraform/Ansible operational model - full rebuilds or new region/provider rollouts in under 1 hour.
- Built multi-site failover and disaster-recovery architecture hands-on: Verizon's TCCP, a multi-site Active-Active container platform on an eBGP Clos fabric with Consul DNS multi-site failover, Docker Swarm, and Ceph storage - sustained multiple concurrent hardware/software failures at 100% uptime over 1.5+ years.
- Shipped self-hosted/BYOC deployment automation for environments customers control: an open-source Ansible Galaxy Collection partners use to launch self-hosted Cube Guardian instances (zero-touch HA HashiCorp Vault provisioning, published deployment docs for self-service setup), plus automation pipelines for deploying and managing enterprise infrastructure in customer environments at Dell.
- Runs infrastructure as code end to end: Terraform/Ansible zero-touch provisioning of a high-security exchange production environment built from scratch (HA PostgreSQL and HashiCorp Vault clusters, KVM hosts), git-backed inventory as the single source of truth (CMDB/IPAM), and GitOps-driven provisioning with GitHub Actions PR preview environments and GitLab-CI pipelines.
- Designed security-critical infrastructure: high-security production environment for a hybrid decentralized crypto exchange, HashiCorp Vault secrets management, co-founded Figment's Sensitive Operations team for sensitive infrastructure, and led security incident discovery, remediation, and post-mortems with SecOps.
- Built observability for distributed fleets from the ground up - Prometheus/Grafana fleet-health dashboards and alerting, Datadog/Vector pipelines, and large-scale Splunk/Dynatrace centralized logging and APM - cutting MTTR from 4+ hours to under 15 minutes - plus OpenTelemetry/Langfuse agent tracing in a hands-on AI-platform lab.
Hands-on senior infrastructure/platform engineer with 25+ years building production systems from zero at fast-moving, security-critical fintech and crypto companies, including a low-latency crypto exchange and 1,000+-server blockchain fleets, with early-career grounding in PCI/SOX-audited card-payment environments. Titles have ranged from IC to Head/Director, but the work stayed hands-on: personally designing, building, and operating the platform. Full-stack depth from bare metal through AWS, Kubernetes, Terraform/Ansible automation, CI/CD, PostgreSQL, and observability, building the platform tooling that lets engineering teams move fast without breaking production.
- Designed and built the high-security production environment from scratch for Cube.Exchange, a low-latency hybrid decentralized crypto exchange: Terraform/Ansible zero-touch provisioning, HA PostgreSQL and HashiCorp Vault clusters, KVM virtualization, blockchain RPC nodes, and Datadog/Vector observability.
- Deep PostgreSQL platform work: developed a Python/Ansible framework for backup, restoration, scaling, tuning, and migrations, designed a 600TB distributed PostgreSQL (Citus) hot store with 11PB cold data in S3 for a telemetry platform built for 1.5M TPS and 120TB/day ingest, and prototyped an AWS OpenSearch-backed analytics assistant.
- Designed standardized Terraform/Ansible platform abstractions (structured inventories, reusable roles, runbooks) to scale infrastructure through hypergrowth: 125+ blockchain networks across 1,000+ servers and six cloud/bare-metal providers, with new-environment provisioning cut from 15+ days to under 24 hours.
- Built CI/CD engineering teams ship on: GitHub Actions pipelines with PR preview environments validating deployments before production, autoscaling self-hosted Kubernetes (EKS) CI runners, and GitLab-CI-driven infrastructure provisioning that cut deploy cycles from 30 days to 8 hours.
- Architected a multi-site Active-Active container platform with eBGP Clos networking, Consul DNS service discovery, Docker Swarm, and Ceph storage that held 100% uptime over 1.5+ years while absorbing concurrent hardware and software failures.
- Built observability and incident response that cut fleet MTTR from 4+ hours to under 15 minutes: Datadog, Prometheus/Grafana fleet dashboards shared across SRE and support, OpenTelemetry/Loki/Tempo tracing, and standardized runbooks with AI-assisted operations.
Hands-on SRE and platform engineer with 25+ years designing, automating, and operating mission-critical, highly available production infrastructure - Linux, Kubernetes, Helm, Terraform/Ansible, and PostgreSQL - across fleets of 1,000+ servers where downtime carries real business and customer consequences. Builds production environments from scratch, owns systems end to end, and treats automation as leverage: zero-touch provisioning, CI/CD with PR preview environments, and observability that cuts MTTR from 4+ hours to under 15 minutes.
- Owned production environments end to end: designed and built Cube.Exchange's high-security, low-latency production environment from scratch - Terraform/Ansible zero-touch provisioning, HA PostgreSQL and HashiCorp Vault clusters, KVM hosts, distributed blockchain RPC infrastructure, and Datadog/Vector observability.
- Architected and operated mission-critical, high-availability distributed systems at 1,000+ server scale: multi-site Active-Active container platform with 100% uptime over 1.5+ years; 99.9% uptime targets on nationwide streaming infrastructure.
- Designed a standardized Ansible/Terraform platform operating 125+ blockchain networks across six cloud and bare-metal providers.
- Deep hands-on production experience with Kubernetes, Helm, Terraform, Ansible, Docker, Linux, PostgreSQL, HashiCorp Vault, GitHub Actions and GitLab CI/CD, Python/Bash automation, Prometheus/Grafana, Datadog, OpenTelemetry, and production incident response.
- Built and optimized deployment pipelines end to end: GitHub Actions CI/CD with PR preview environments validating sync, OS, and network metrics before production, autoscaling Kubernetes CI runners (GitHub Actions Runner Controller), Rust build/test pipelines, and GitLab-CI-driven infrastructure deployment - cutting new-environment provisioning from 15+ days to under 24 hours and streamlining developer workflows (DevX).
- Ran Kubernetes in production for CI infrastructure and maintains a production-shaped k3s platform lab - Cilium ingress, Helm/Kustomize delivery, secure service exposure, Prometheus/Grafana, OpenTelemetry, Loki/Tempo - including a LiteLLM model-serving gateway (Python) with caching, guardrails, and budget controls.
- Automation as leverage and standards for operational excellence: Python/Ansible framework for PostgreSQL backup, restore, scaling, and migrations; runbooks and fleet-wide upgrade workflows; AI-assisted engineering standards that cut MTTR from 4+ hours to under 15 minutes.
Hands-on infrastructure lead with 25+ years designing, operating, and troubleshooting highly available cloud, hybrid, and data-center infrastructure across fintech, telecom, and blockchain. Standardizes heterogeneous providers onto one Terraform/Ansible operational model: ran 1,000+ server fleets across AWS and five bare-metal/COLO providers (DigitalOcean, Latitude, OVH, Servers.com, Nirvana Cloud) in 13+ regions (Nirvana Labs) and 1,300+ servers across 13 cloud/COLO providers (Figment), and operated infrastructure across 75+ data centers for a nationwide OTT/IPTV platform (Verizon OnCue). Owns reliability end to end: a multi-site Active-Active platform that held 100% uptime over 1.5+ years, fleet MTTR driven from 4+ hours to under 15 minutes, and 1st/2nd/3rd-level escalation for critical production incidents, running AWS, hybrid cloud, IaC, observability, and cost optimization hands-on, both as direct infrastructure owner and, at Dell, as delivery lead for external enterprise and service-provider customers.
- Standardizes heterogeneous providers onto one operational model: unified fleet operations (provisioning, upgrades, telemetry ingestion, TLS/certificate and firewall workflows) with Terraform and Ansible across six cloud and bare-metal providers (AWS, DigitalOcean, Latitude, OVH, Servers.com, Nirvana Cloud) and 13+ regions, giving the team one model regardless of provider, with new region/provider rollouts in under 1 hour.
- Owns availability and reliability hands-on: designed a multi-site Active-Active container platform (eBGP Clos fabric, Consul DNS multi-site failover, Docker Swarm, Ceph storage) that sustained concurrent hardware and software failures at 100% uptime over 1.5+ years, and built fleet observability (Datadog, Prometheus/Grafana) that drove MTTR from 4+ hours to under 15 minutes.
- Leads cloud migrations, modernization, and data-center integration: sequenced Verizon's analytics platform cutover from AWS-managed services to self-hosted OpenShift to cut opex and vendor lock-in, and migrated 30+ critical systems to a new UK data center with zero downtime at TSYS.
- Acts as hands-on escalation and troubleshooting point for complex compute, storage, networking, and connectivity issues: maintained 24x7 on-call with direct client interaction and 1st/2nd/3rd-level escalation for critical incidents at TSYS (a credit-card processor), where a compliance-remediation platform held a 100% audit pass rate across 7 network environments (SAS 70, PCI DSS, SOX 404), and led security-incident discovery, remediation, and post-mortems with SecOps at Figment.
- Designs and operates infrastructure at data-center scale and cost: architected a nationwide OTT/IPTV platform (1,000+ edge locations, 40+ POPs, 20+ content data centers, 15+ Pbps) with infrastructure operated across 75+ data centers, and drove server, storage, and network vendor evaluations delivering $80M and $10M+ in cost savings.
- Builds secure, resilient production environments from scratch: Terraform/Ansible zero-touch provisioning of HA PostgreSQL and HashiCorp Vault clusters, KVM hosts, and blockchain RPC nodes for a low-latency crypto exchange (Cube.Exchange), with Datadog/Vector observability and GitHub Actions CI/CD.
Infrastructure leader with 25+ years running mission-critical Linux infrastructure at telecom scale and building the 24/7 teams behind it: founded the first DevOps team in Verizon Telematics, led infrastructure and automation teams at Dell, Figment, and Nirvana Labs, and owned $200M+ CAPEX across 75+ data centers. Brings the hardware-layer depth to lead a GPU cluster team credibly, plus the ticketing, on-call, and vendor discipline that uptime-driven customer infrastructure demands.
- Builds and leads 24/7 infrastructure teams: a global node-operations team of 5 at Nirvana Labs (hired 3 seniors; established onboarding, on-call, and code-review standards), 10 hires and 5 promotions at Figment, and 8+ engineers and contractors on Verizon Telematics' first DevOps team.
- Defined the SLA (99.999% uptime) for TCCP, a multi-site Active-Active platform on high-density OCP racks with eBGP Clos networking and Ceph storage, then exceeded it with 100% uptime over 1.5+ years; also sustained 99.9% uptime on a nationwide OTT video platform.
- Brings hands-on GPU and server-hardware depth: benchmarked and performance-tuned a Dell C410x lab with 16x NVIDIA GPUs, resolving a PCIe over-subscription bottleneck (supporting $9M+ in encoding cost savings), and authored an open-source Terraform provider for Dell PowerEdge firmware/BIOS lifecycle via Redfish/BMC APIs.
- Runs Linux fleets as multi-tenant infrastructure: 1,000+ servers across six providers at Nirvana Labs (125+ blockchain networks as isolated, always-on tenant workloads) and 1,300+ across 13 providers at Figment on one Terraform/Ansible model, with observability and AI-assisted runbooks that together cut fleet MTTR from 4+ hours to under 15 minutes.
- Owns vendor and customer relationships day to day: 24x7 on-call with direct customer escalations at TSYS, customer-facing solution architecture and RFI/RFQ ownership as Director of DevOps at Dell, vendor negotiations delivering $80M and $10M+ in savings at Verizon, and ticketing/incident management with Linear, JIRA, GitLab, PagerDuty, and Incident.io.
Hands-on site reliability and infrastructure engineer with 25+ years building highly reliable infrastructure at scale across fintech, telecom, and blockchain industries, from large multi-chain node fleets (125+ blockchain networks, 13+ VM/BM providers) to low latency/high security crypto exchanges, and multi-site active-active container platforms, all built on Terraform/Ansible/Kubernetes & GitOps.
- Designed and deployed fully automated blockchain node fleets for over 125 blockchains, including Arbitrum One (Classic & Nitro) Full/Archive/Relay nodes, Arbitrum L3 chains (Xai/Corn/ApeChain), and multiple OP Stack chains and Ethereum EL/CL clients (geth/reth/erigon/Lighthouse/Prysm), across over 13 multi-region cloud & bare-metal providers (+COLO).
- Designed the standardized Terraform/Ansible platform (structured inventories, reusable roles, runbooks), cutting new-network provisioning from 15+ days to under 24 hours and enabling full rebuilds or new region/provider rollouts in under 1 hour.
- Built end-to-end GitHub Actions CI/CD with PR preview environments that validate node sync, OS/network metrics, and dashboards before production, plus autoscaling self-hosted CI runners on Kubernetes with GitHub Actions Runner Controller.
- Built observability and reliability engineering (Prometheus/Grafana/Loki, Datadog/Vector, Splunk) with dashboards, alerting, incident response, and postmortems that cut org MTTR from 4+ hours to under 15 minutes.
- Diagnose low-level networking and storage issues across distributed systems: ZFS snapshot/restore with NVMe 4K block alignment for fast node recovery, eBGP Clos fabrics with Ceph storage (100% uptime over 1.5+ years), and secure-by-default production environments built from scratch with zero-touch Terraform/Ansible and HashiCorp Vault.
Hands-on engineering leader with 25+ years building and scaling both the engineering teams and the platform, CI/CD, and DevOps/SRE foundations they run on across telecom, fintech/crypto, and blockchain, owning vision, roadmap, and execution while treating developer velocity and platform reliability as the product and engineers as the customer. Now driving AI-native developer productivity (reusable agents, MCP, AI-assisted PR review, LLM tooling) into engineering workflows in high-stakes, high-reliability environments.
- Led globally distributed engineering, DevOps, and platform organizations across telecom, fintech/crypto, and blockchain: a 5-person node DevOps team operating 125+ blockchain networks on 1,000+ servers across 13+ regions at Nirvana Labs, the first DevOps team in Verizon Telematics (8+ engineers), and Verizon's Product Analytics Cloud software/data engineering team (20M routers, 1.5M TPS, 120TB/day).
- Drove AI-native engineering and developer productivity: established reusable agent standards, MCP integrations, agent tracing/observability, and AI-assisted PR reviews that drove 5x faster new-network role development and cut MTTR from 4+ hours to under 15 minutes, and prototyped an LLM text-to-SQL assistant to demonstrate AI-native self-serve analytics on a production platform.
- Built end-to-end engineering platforms and paved roads: a standardized Ansible/Terraform platform (reusable roles, runbooks) that deployed 125+ networks and cut new-network provisioning from 15+ days to under 24 hours, GitHub Actions and GitLab CI/CD with PR preview environments, and autoscaling self-hosted Kubernetes/EKS runners for pipeline parallelism and isolation.
- Built a production-shaped AI platform lab: a LiteLLM OpenAI-compatible gateway with caching, guardrails, and budget controls, a LangGraph docs RAG agent, a four-backend vector bake-off, and an LGTM observability/eval stack (Grafana, Loki, Tempo, Prometheus, Langfuse) on Kubernetes with OpenTelemetry tracing.
- Delivered reliability as a contract in high-stakes systems: the TCCP multi-site Active-Active container platform held 100% uptime over 1.5+ years and cut deploy cycles from 30 days to 8 hours; built Cube.Exchange's high-security production environment for a low-latency crypto exchange with Terraform/Ansible zero-touch provisioning and HashiCorp Vault, and led security incident response and post-mortems with SecOps at Figment.
- Hired, coached, and grew engineering teams through rapid scaling (10 hires / 5 promotions at Figment, 3 senior engineers at Nirvana Labs) while partnering with product, security, finance, and executive stakeholders and aligning technical decisions with delivery milestones in remote-first, documentation-heavy cultures.
Hands-on engineering manager who stands up the secrets and security infrastructure an organization does not have yet and makes it pass audit: HashiCorp Vault HA clusters as a production secrets management backend, an open-sourced Vault provisioning collection, and TLS certificate and firewall workflows automated across a 1,000+ server AWS and bare-metal fleet. Security is the foundation, not a pivot: an A.A.S. in Cyber Defense earned alongside a cardholder-data-protection and compliance platform at a credit card processor in the early 2000s (Visa CISP/PCI DSS, SOX 404, CIS/NIST hardening). 25+ years since across telecom, payments/fintech, and blockchain, repeatedly entering unfamiliar domains and reaching operating depth fast.
- Supervised engineers for 8+ years across six engineering leadership roles, including: led a senior Verizon software and data engineering team with two Distinguished Engineers (DMTS) plus contractor teams; hired 10 engineers and coached 5 to promotion at Figment; founded the first DevOps team in Verizon Telematics (8+ employees and contractors); and led a 5-person global node DevOps team at Nirvana Labs, hiring 3 senior engineers and setting onboarding, on-call, and code-review standards.
- Built and operated HashiCorp Vault hands-on as Head of DevOps at Cube.Exchange: Ansible roles and playbooks that initialize and manage HA Vault clusters as the secrets management backend for a high-security production environment, an open-sourced Ansible Galaxy Collection that zero-touch provisions HA Vault clusters for partner self-hosted deployments, and TLS certificate, load balancer, and firewall workflows automated with Terraform across a 1,000+ server fleet.
- Owned security-critical work in regulated environments: co-founded Figment's Sensitive Operations team to harden sensitive infrastructure and validator onboarding, led security incident discovery, assessment, remediation, and post-mortems with SecOps, and designed a security-compliance and cardholder-data-protection platform at a credit card processor (SAS 70, Visa CISP/PCI DSS, SOX 404, CIS/NIST hardening) sustaining a 100% audit pass rate across 7 network environments in the US and UK.
- Delivered large-scale, mission-critical platforms end to end, sequencing migrations and dependencies with stakeholders across support, SRE, leadership, and account teams: TCCP, a multi-site Active-Active container platform, held 100% uptime over 1.5+ years and cut deploy cycles from 30 days to 8 hours; a standardized Ansible/Terraform platform deployed 125+ networks across 1,000+ servers, 13+ regions, and six cloud and bare-metal providers including AWS, cutting provisioning from 15+ days to under 24 hours.
- Ran monitoring and metrics as the input to decisions: Splunk and Dynatrace AppMon application performance monitoring, Datadog/Vector pipelines, and Grafana dashboards shared with support, SRE, leadership, and account teams, plus Python capacity analytics and automated cost reporting that drove consolidation roadmaps; instrumentation and triage standards cut MTTR from 4+ hours to under 15 minutes.
- Coded and shipped in Python, Go, and Bash with open-source and REST API integration depth: authored an open-source Go Terraform provider driving server configuration and firmware lifecycle through Redfish/BMC REST APIs, published an Ansible Galaxy Collection, served on the OpenSwitch TSC under the Linux Foundation, and built CI/CD that gates changes on unit tests, static analysis, and PR preview environments before production.
Hands-on platform, DevOps, and site reliability (SRE) engineer with 20+ years building and running shared infrastructure platforms across AWS, Kubernetes-based environments (EKS, OpenShift, k3s), and traditional data centers, including PCI/SOX-audited card-payment platforms and a low-latency crypto exchange. Personally designs, builds, and operates the Terraform/Ansible Infrastructure as Code (IaC), configuration management, CI/CD, and monitoring/observability that engineering teams ship on; titles have ranged from IC to Head/Director, but the work stayed hands-on. Python-first automation with Go for systems tooling, and a track record of turning manual operations into reliable, scalable, self-service platforms.
- Designed the shared Terraform/Ansible platform a global node-operations team uses to deploy and operate 1,000+ servers across six cloud and bare-metal providers: new-environment provisioning cut from 15+ days to under 24 hours, full rebuilds in under 1 hour.
- Built the CI/CD engineering teams ship on: GitHub Actions pipelines with PR preview environments validating deployments before production, autoscaling self-hosted Kubernetes (EKS) CI runners, and GitLab-CI-driven infrastructure provisioning that cut deploy cycles from 30 days to 8 hours.
- Built observability and incident response that cut fleet MTTR from 4+ hours to under 15 minutes: Datadog and Vector pipelines, Prometheus/Grafana dashboards shared across SRE and support teams, Splunk and Dynatrace log/APM platforms, and standardized runbooks, postmortems, and AI-assisted operations.
- Designed and deployed a high-security production environment from scratch for Cube.Exchange, a low-latency hybrid decentralized crypto exchange: zero-touch Terraform/Ansible provisioning, HA PostgreSQL and HashiCorp Vault clusters, KVM virtualization, blockchain RPC nodes, and Datadog/Vector observability.
- Develop software and automation in Python, Bash, and Go: Python/Ansible frameworks for PostgreSQL backup, scaling, and migrations, Python capacity analytics for a 1,300+ server fleet, and an open-source Go Terraform provider for Dell PowerEdge servers driving configuration and firmware lifecycle through Redfish/BMC REST APIs.
- Operate hybrid infrastructure at high stakes: AWS plus COLO and traditional data centers, containerized workloads from Docker Swarm to EKS, OpenShift, and k3s, a zero-downtime migration of 30+ critical systems in PCI/SOX-audited card-payment environments, and 24x7 on-call as 1st through 3rd level escalation.
Hands-on platform and infrastructure engineer with 25+ years building and operating the AWS, Kubernetes, and bare-metal systems other engineers ship on: the Terraform, Ansible, and Python automation, the GitHub Actions and GitLab CI/CD pipelines, and the Grafana and Datadog observability behind production fleets of 1,000+ servers across six cloud and bare-metal providers. Recent work runs toward ML and LLM infrastructure: a self-built Kubernetes (k3s) AI platform lab with a LiteLLM model-serving gateway, Langfuse and LGTM observability (Prometheus, Loki, Tempo, OpenTelemetry), and FastAPI/RAG services, plus petabyte-scale telemetry and analytics data platforms.
- Designed the standardized Terraform and Ansible automation (structured inventories, reusable roles, runbooks) that deploys and operates 1,000+ servers across six cloud and bare-metal providers including AWS, cutting new-environment provisioning from 15+ days to under 24 hours and putting fleet upgrades, telemetry ingestion, load balancer/TLS, and firewall workflows behind one operational model.
- Deployed and operated containerized workloads in production: autoscaling self-hosted CI runners on Kubernetes/EKS via GitHub Actions Runner Controller, the planning and cutover sequencing that moved analytics platform components from AWS-managed services to self-hosted OpenShift, and a multi-site Active-Active Docker Swarm/Ceph container platform that held 100% uptime over 1.5+ years while cutting deploy cycles from 30 days to 8 hours.
- Developed the automation and tooling directly: a Python/Ansible framework for PostgreSQL backup, restore, scaling, tuning, and migrations, Python capacity analytics across a 1,300+ server fleet, an open-source Go Terraform provider driving Dell PowerEdge configuration and firmware lifecycle through Redfish/BMC REST APIs, and GitHub Actions pipelines with PR preview environments, unit testing, and static analysis.
- Built the model-serving and observability layer of a Kubernetes (k3s) AI platform lab: LiteLLM as an OpenAI-compatible gateway for local and hosted models with virtual keys, caching, guardrails, and budget controls, Langfuse plus an LGTM stack (Grafana, Loki, Tempo, Prometheus, Alloy, OpenTelemetry Collector) for traces, evals, dashboards, logs, and metrics, and FastAPI/RAG services instrumented with Prometheus metrics and OTLP tracing.
- Raised production operability on live fleets: AI standards with reusable agents, skills, MCP integrations, and agent tracing cut MTTR from 4+ hours to under 15 minutes and made new automation-role development 5x faster, while Datadog/Vector pipelines, shared Grafana dashboards, and Splunk/Dynatrace APM gave support, SRE, and engineering teams a common view of the systems they run.
Infrastructure and telemetry engineer with 25+ years inside large datacenter and telecom environments: eBGP Clos fabrics and multi-site container platforms, fleets of 1,000+ servers across 13+ regions and six cloud and bare-metal providers, Dell server configuration and firmware automation through Redfish/BMC REST APIs, and router telemetry platforms engineered for 20M routers and 120TB/day ingest. Now writing the agent layer on top of that operational substrate: team-wide AI engineering standards with reusable agents, Model Context Protocol (MCP) integrations, and agent tracing adopted in daily operations at Nirvana Labs, plus a self-built, open-source Docker and Kubernetes AI platform lab with a FastMCP tool server, a LangGraph RAG agent served over FastAPI, a LiteLLM gateway, and Prometheus, Loki, Grafana, and OpenTelemetry instrumentation.
- Established the AI engineering standards at Nirvana Labs: reusable agents, skills, Model Context Protocol (MCP) integrations, observability and agent tracing, and AI-assisted PR reviews, adopted across the node DevOps team operating 125+ blockchain networks on 1,000+ servers in 13+ regions, driving 5x faster new-network role development and cutting MTTR from 4+ hours to under 15 minutes on a fleet where standardized automation took new-network provisioning from 15+ days to under 24 hours.
- Built the agent, tool, and retrieval layers hands-on in Python in a self-built, open-source Kubernetes (k3s) AI platform lab: a FastMCP retrieval server exposing normalized search, document and chunk lookup, corpus stats, and cross-backend comparison as tools; a LangGraph docs RAG agent with deduplication and reranking, LiteLLM synthesis, and FastAPI serving; and a four-backend vector bake-off (Qdrant, Weaviate, pgvector, Milvus) with shared ingestion, chunking, embedding, and search code. Also prototyped an LLM text-to-SQL analytics assistant over schema and telemetry-metric context in AWS OpenSearch vector search on Verizon's router telemetry platform.
- Wired the telemetry and log pipelines operators actually run on: Datadog with Vector data pipelines at Cube.Exchange, Grafana fleet-health dashboards shared with support, SRE, leadership, and account teams at Nirvana Labs, large-scale Splunk log aggregation and Dynatrace AppMon at Verizon Connect, and DataDog time-series dashboards for infrastructure metrics and application logs at Figment, plus a Langfuse and LGTM stack (Grafana, Loki, Tempo, Prometheus, Alloy, OpenTelemetry Collector) in my personal AI platform lab for traces, evals, dashboards, logs, and metrics.
- Automated servers and network fabric at the API and protocol layer, the same surface an infrastructure agent has to reach: an open-source Go Terraform provider driving Dell PowerEdge configuration and BIOS/RAID firmware lifecycle through Redfish/BMC REST APIs, published publicly and adopted across Dell's IaaS, Hybrid Cloud, and Telco/Service Provider teams; eBGP Clos fabric design and multi-site operation; and a Linux Foundation OpenSwitch Technical Steering Committee seat representing Verizon Connect on open network OS direction.
- Built datacenter and network infrastructure at telecom scale: designed and built TCCP, a multi-site Active-Active container platform on high-density OCP racks with an eBGP Clos fabric, Consul DNS service discovery, Docker Swarm, and Ceph, holding 100% uptime over 1.5+ years and cutting deploy cycles from 30 days to 8 hours; and architected a nationwide OTT and IPTV streaming platform spanning 1,000+ edge locations, 40+ POPs, and 15+ Pbps of throughput.
- Wrote the automation and analysis code directly: a Python/Ansible framework for PostgreSQL backup, restore, scaling, tuning, and migrations at Cube.Exchange, Python spare-capacity analytics and a consolidation roadmap on Figment's 1,300+ server fleet, exploratory ML models for router telemetry analysis plus Jupyter, Pandas, and Matplotlib notebooks over Redshift telemetry data at Verizon, and GitHub Actions pipelines with PR preview environments validating sync and OS/network metrics at Nirvana Labs alongside unit testing and static analysis gating Rust services at Cube.Exchange.
- Built toward auditability, reproducibility, and reversibility from the start: a security-compliance analysis and remediation platform at a credit card processor (SAS 70, Visa CISP/PCI DSS, SOX 404, DOED OIG) sustaining a 100% audit pass rate across 7 network environments in the US and UK, zero-touch patch deployment with automated completion verification and failure alerting that reclaimed 500+ engineer-hours annually, fleet inventory kept as a version-controlled git-backed Ansible CMDB and IPAM source of truth, ZFS snapshot/restore for fast node recovery, and DevOps best practices advised and mentored across internal teams and customers at Dell.
Infrastructure engineering leader who builds the teams, operating rhythm, and technical direction behind large-scale bare-metal and multi-cloud fleets. Founded the first DevOps team in Verizon Telematics, hired 10 engineers and coached 5 to promotion at Figment, and led a global node DevOps team at Nirvana Labs that cut fleet MTTR from 4+ hours to under 15 minutes. 25+ years across Site Reliability Engineering (SRE), network fabric, and distributed storage, from a multi-site Active-Active internal cloud on high-density OCP racks with an eBGP Clos fabric and Ceph (100% uptime over 1.5+ years) to a $200M+ CAPEX budget across 75+ data centers. Hardware depth reaches the accelerator layer: benchmarked and performance-tuned a Dell C410x lab with 16x NVIDIA GPUs, diagnosing a PCIe over-subscription bottleneck and re-architecting to 8 GPUs/server.
- Builds and leads distributed infrastructure teams across five orgs: founded the first DevOps team in Verizon Telematics (8+ employees and contractors), hired 10 engineers and coached 5 to promotion at Figment, led a customer-facing DevOps team as Director at Dell Technologies, led a senior software and data engineering team as Associate Director at Verizon including two Distinguished Engineers (one leading his own contractor team), and ran a global node DevOps team of 5 at Nirvana Labs (hired 3 senior engineers and established documentation, onboarding, on-call, and code-review standards).
- Drives Site Reliability Engineering (SRE) and delivery metrics that move: 100% uptime over 1.5+ years on TCCP and 99.9% on a nationwide OTT video platform, MTTR from 4+ hours to under 15 minutes, new-network provisioning from 15+ days to under 24 hours, fleet-wide upgrades and full region/provider rebuilds in under 1 hour, and datacenter rebuild and deploy cycles from 30 days to 8 hours, backed by Datadog, Prometheus/Grafana, and Splunk observability plus structured runbooks, incident response, on-call standards, and post-mortems.
- Architects production network fabrics and nationwide edge footprints: TCCP, a multi-site Active-Active internal cloud on high-density OCP racks built on an eBGP Clos (spine-leaf) fabric with Consul DNS service discovery and multi-site failover; and Verizon OnCue OTT/IPTV across 1,000+ edge locations, 40+ POPs, and 20+ content ingest/encoding data centers at 15+ Pbps.
- Has built and operated distributed storage across block, file, and object, tuned for throughput as well as capacity: Ceph clusters behind TCCP with Ansible-automated cluster deployment, ZFS snapshot/restore with NVMe 4K block-size alignment for high-throughput disk I/O and fast node recovery across a 1,000+ server fleet, a 1.5PB Hadoop/HDFS cluster, and a hot/cold design of 600TB compressed Postgres/Citus plus 11PB in Dremio over S3 for a telemetry platform sized for 120TB/day ingest.
- Standardizes bare-metal and multi-cloud fleets on one Terraform/Ansible operational model: 1,000+ servers across six cloud and bare-metal providers at Nirvana Labs and 1,300+ servers across 13 cloud/COLO providers at Figment, with GitHub Actions CI/CD and PR preview environments, Kubernetes/EKS autoscaling runners, and an open-source Go Terraform provider automating Dell PowerEdge firmware/BIOS lifecycle through Redfish/BMC APIs.
- Converts scale into budget and roadmap: owned planning and tracking of a $200M+ CAPEX budget across 75+ data centers, drove server/accelerator/GPU/encoder evaluation and vendor negotiations delivering $80M and $10M+ in savings, led an AWS-managed-services to self-hosted OpenShift migration to cut opex and vendor lock-in, and built Python capacity analytics that drove an infrastructure consolidation roadmap.
Hands-on infrastructure engineer with 25+ years owning compute, networking, storage, and colocation end to end, with a track record as the primary customer technical contact for strategic accounts. At Dell, owned solution design, RFI/RFQ responses, and issue resolution for strategic Enterprise, Telco, and Service Provider accounts in EMEA; at TSYS, a credit card processor, held 24x7 on-call with direct client interaction as 1st, 2nd, and 3rd level escalation. Operates at fleet scale: 1,000+ servers across six cloud and bare-metal providers at Nirvana Labs, 1,300+ servers across 13 cloud/COLO providers at Figment, and a multi-site Active-Active platform on high-density OCP racks with an eBGP Clos Ethernet fabric and Ceph distributed storage that held 100% uptime over 1.5+ years, with fleet MTTR cut from 4+ hours to under 15 minutes.
- Owns compute across cloud, bare-metal, and colocation providers: standardized fleet operations (provisioning, upgrades, telemetry ingestion, load balancer/TLS certificate and firewall workflows) on one Terraform/Ansible model across six providers (AWS, DigitalOcean, Latitude, OVH, Servers.com, Nirvana Cloud) covering 1,000+ servers at Nirvana Labs, plus 1,300+ servers across 13 cloud/COLO providers at Figment, cutting new blockchain-network provisioning from 15+ days to under 24 hours and new region/provider rollouts to under 1 hour.
- Primary customer technical contact for strategic accounts: owned solution design, RFI/RFQ responses, and issue resolution at Dell, presenting infrastructure solutions and business cases directly to enterprise and Telco customers, and drove customer technical discussions to capture requirements, architect enterprise infrastructure solutions, and build the automation pipelines that ran in those customer environments.
- Owns observability and issue lifecycle as one loop: built Grafana fleet-health dashboards giving support, SRE, leadership, and account teams shared visibility into fleet and node operations, ran Datadog/Vector pipelines plus large-scale Splunk and Dynatrace AppMon deployments, and led security incident discovery, assessment, remediation, and post-mortem analysis with SecOps at Figment.
- Runs networking and storage at the hardware layer: designed a multi-site Active-Active platform on high-density OCP racks with an eBGP Clos Ethernet fabric, Consul DNS multi-site failover, Docker Swarm, and Ceph distributed block/file/object storage, sustaining concurrent hardware and software failures at 100% uptime over 1.5+ years; and architected a nationwide OTT/IPTV platform at 15+ Pbps across 1,000+ edge locations, 40+ POPs, and 20+ content data centers.
- Carries escalation and multi-data-center delivery: maintained 24x7 on-call with direct client interaction as 1st, 2nd, and 3rd level escalation for critical issues at TSYS, a credit card processor, migrated 30+ critical systems to a new UK data center with zero downtime, and designed and deployed server, storage, and network infrastructure across 75+ data centers for go90 and FiOS IPTV.
- Reaches the hardware and firmware layer in code: authored an open-source Go Terraform provider for Dell PowerEdge servers automating bare-metal configuration and firmware/BIOS lifecycle (BIOS settings, BIOS and RAID-controller firmware updates) through Redfish/BMC REST APIs, and benchmarked a Dell C410x GPU video-encoding lab (16x NVIDIA GPUs), diagnosing a PCIe over-subscription bottleneck and re-architecting across two servers at 8 GPUs/server in support of $9M+ in encoding cost savings.