Site Reliability Engineer

I'm Devin, the SRE you hire when things need to stay up and costs need to come down.

I help startups and scale-ups ship reliably, cut cloud costs, and get their infrastructure under control. Senior SRE for contract work — direct access, no account managers.

Devin Collins
19 years in the trenches
100+ AWS accounts as code
$500k saved per year on telemetry
100s of pipelines built

/ About me

I've been keeping systems up since 2007

I'm a Principal Site Reliability Engineer based in Southern California, working remotely with clients around the world. For nearly 20 years I've built and run the platforms behind Geek Squad, Malwarebytes, JupiterOne, and Skillz — the kind of systems that need to stay up while people are counting on them.

My sweet spot is the messy middle: the company that has outgrown its infrastructure, or the platform team that needs senior hands without a full-time hire. I bring Kubernetes, Terraform, AWS, observability, and incident response experience and apply it pragmatically. I care about systems that are boring to run, because boring is what makes them reliable.

Previously worked on
Malwarebytes JupiterOne Skillz Best Buy / Geek Squad

/ My skills

What I'm strong at

If your problem lives in one of these buckets, there's a good chance I can help.

Kubernetes

Cluster rollout, adoption, and day-2 operations that don't turn into a second job.

Terraform

Infrastructure as code from a single account up to 100+ AWS accounts managed as one codebase.

AWS

Well-Architected reviews that find real problems and real savings, often at the same time.

Observability

Monitoring, alerting, and SLOs teams actually trust. I cut MTTR and the telemetry bill at once.

CI/CD

GitHub Actions, Jenkins, ArgoCD — releases should be predictable and boring.

Incident Response

Contract on-call, alerting that pages for the right reasons, and blameless postmortems.

Linux

Two decades of the good stuff — the operating system underneath nearly everything I run.

Go & Automation

Tooling that removes the toil. If a task is worth doing twice, I automate it.

/ Selected work

Things I've built and run

Client work is confidential, but a few things I've done speak for themselves.

ObservabilityAWS

Observability stack migration

Moved a mixed self-hosted/CloudWatch stack to New Relic at JupiterOne. Better insight for every team, lower MTTR, an MTTD baseline — and more than $500k saved per year.

KubernetesPlatform

Kubernetes rollout & adoption

Took JupiterOne from first clusters to business-as-usual, simplifying deploys across internal applications and setting standards so new services shipped with observability and testing built in.

TerraformAWS

100+ AWS accounts as code

Managed 100+ AWS accounts from a single Terraform codebase — reproducible infrastructure, reviewable changes, no drift. Written up here.

MobileScale

MyTLC

A mobile application I built that supported 175k+ active users, syncing daily work schedules to their phones. Lessons from that scale inform how I design infrastructure today.

/ Services

Ways I can help

Whether it's an ad-hoc request or a longer engagement. If you're not sure something fits, just ask.

Kubernetes Adoption & Operations

Cluster design, rollout, and ongoing operations. Adopt Kubernetes without adopting the chaos that usually comes with it.

Terraform & Multi-Account AWS

Infrastructure as code at any scale, including 100+ accounts from one codebase. Standardized, reviewable, reproducible.

Observability Stack Migration

From a mess of self-hosted tools and CloudWatch to a coherent stack — with the MTTR and bill both going down.

CI/CD Pipelines

Faster, safer build, test, and deploy pipelines across GitHub Actions, Jenkins, and ArgoCD.

AWS Well-Architected Review

A structured review for performance, scalability, security, and cost, with recommendations you can actually act on.

Cloud Cost Optimization

Find wasted resources, right-size what's over-provisioned, and get predictable spend without sacrificing reliability.

Incident Response & On-Call

Take over or augment on-call, or build the alerting, runbooks, and process so your own team can handle incidents.

SLOs, SLIs & Error Budgets

Define the numbers that matter, wire up measurement, and turn reliability into something you can reason about.

Cloud Migration & Modernization

Move workloads to the cloud or into containers in staged, reversible steps. No big-bang cutovers.

Security & Compliance Readiness

Practical security improvements and SOC 2-style audit readiness. Secure enough, without strangling the team.

Performance & Capacity Tuning

Find bottlenecks, tune queries, and optimize code so systems keep up with your users and your growth.

DevOps Guidance

Work alongside your team to improve delivery, reliability, and confidence. Without the cargo cult.

/ FAQ

Questions I get asked a lot

Who are you?

I'm Devin, a Principal Site Reliability Engineer in Southern California. After nearly 20 years in the trenches — from Geek Squad's service desk to Principal SRE at Malwarebytes and JupiterOne — I now do contract work for companies that need senior infrastructure help without a full-time hire.

Why should we hire you?

You get a senior generalist who has seen a lot of failure modes and knows how to avoid them. I deliver work at a high standard, with solutions that work now and in the future. If you have small-to-medium projects and no spare capacity, I temporarily boost your team without the overhead of a hire.

What's your experience?

Principal SRE at Malwarebytes, Senior SRE at JupiterOne, Lead Software Engineer at Skillz before that. I've rolled out Kubernetes, moved organizations onto Terraform-managed infrastructure, built hundreds of pipelines, set up monitoring and alerting, and kept a heavy focus on cost control and compliance like SOC 2.

How does payment work?

After an initial consultation, if we agree on the work, I'll provide a contract outlining the payment terms. I offer fixed-price and retainer engagements; time-and-materials is available for open-ended work.

For fixed-price agreements, a good-faith deposit is required before work starts, with the balance paid upon delivery of agreed milestones. Rates are agreed up front — ask and we'll figure out what fits your budget.

Do you outsource any work?

No. When you work with me, you work with me.

Do you work in our office?

I'm fully remote and currently only accepting remote projects. I work Pacific time and have shipped with teams in every timezone.

/ Get in touch

Interested in working together? Let's talk.

Send me a message about your project, or email me directly. I reply within 24 hours.