Run operational tasks with background agents

Proactively run tasks like deployment monitoring, operational reports, and resource optimization on a schedule or on trigger

Automations

Automate tasks on schedule or on trigger

Run your workflows on a cadence, or fire agents on event triggers like deploys, alerts, and other production events

Engineering Deploy's Agent

Watches #engineering for deploys; runs post-deploy health checks.

On deploy event

Daily Pulse

Posts daily deploy + incident digest at 7 AM PT.

Daily at 7:00 AM

PostgreSQL Watch

Monitors orders-db-cluster for query regressions and lock contention.

On alert event

Capacity Review

Weekly resource-utilization report across clusters and services.

Weekly · Fridays

Templates · start from a patternBrowse all
Deploy health monitorWatch a channel for deploys; run post-deploy checks.Use template
Daily digestRoll up daily activity from N channels into a morning summary.Use template
Alert triagerWatch an alert channel; classify and route to the right team.Use template
Resource reportPeriodic utilization summary across clusters or services.Use template
Tasks and skills

Configure agents across all your operational work

Tasks like deployment monitoring, operational reports, resource optimization, and more with a template library to start from

WritePreview
Last updated by Maya Lin · 5 days agoHistory
Titledeploy-health
TypeSkillActive
ScopeReads#engineering#engineering-log-prodGitHub ActionsPosts todeploy thread

Deploy Health Skill

Triggered by a deploy notification in #engineering (compare URL, smoke test link, "deploy kicked off"). Post all updates in the deploy thread.

Environments

EnvironmentTypeLog channelCluster
dev0Pre-prod (staging)#engineering-log-devdev0-cluster
app0Production#engineering-log-prodapp0-cluster
rocketProduction (dogfood)#engineering-log-prodapp0-cluster

Environment progression: dev0app0 rocket. Each gates the next.

Conventions

  • Deploy tag: {github_run_id}.{attempt}-{7_char_sha} (e.g. 24893095514.1-dac7506). The github_run_id is the dev0 build run, not the prod deploy run.
  • Tone: Operational. Lead with what broke, what's at risk, what's next.

Pipeline gate

Before running the health check, verify the prod deploy pipeline completed successfully for this deploy tag (Deploy to Cells workflow for app0, Deploy to Rocket from Cells for rocket).

Priority feed

Start your day with a curated priority list

Background agents pre-investigate every item so you open Resolve to verified findings and recommended next steps

Good morning. Overnight, agents reviewed 47 items across 6 channels.

4 need your attention · 2 scheduled for today · 41 resolved cleanly

Production deploy queued behind in-progress run

blockedP1

Finding: Deploy 8ab7b0a hasn't started on app0-cluster — previous deploy still in progress. Health check will run once the rollout begins.

Investigated · 4 min ago

Daily summary skipped — agent stale at run window

blockedP2

Finding: Scheduled 8:00 AM run did not fire — agent was unresponsive between 04:35–16:29 UTC. Will resume on the next scheduled run tomorrow.

Investigated · 11 min ago

checkout-router: p99 latency drifted +172% post-deploy

recheck +15mP2

Finding: Likely candidate: new code path in order-fulfillment.ts hits a slow query at p99 only under load. Re-checking after next traffic window.

Investigated · 38 min ago

auth-gateway rollback after 4m (deploy-time CI signal)

resolved

Finding: Caught at deploy time and rolled back cleanly. No customer impact recorded. No further action needed.

Investigated · 11h ago

Chat

Chat with your agents to dig in

Use chat for live debugging, multi-turn investigations, or to trigger background agents on demand

Wednesday, May 13
VK
You2:14 PM

What's going on with checkout-router p99 today? Is it the deploy at 10:14 PT?

Engineering Deploy's Agent2:14 PM

Yes — drift starts ~2h after the deploy. p99 was steady at ~142ms through 12:00 PT, then climbed to ~387ms as traffic ramped.

Likely candidate: new code path in order-fulfillment.ts hits a slow query at p99 only under load. Error rate and throughput look normal — this is latency-only.

VK
You2:16 PM

Set up an agent to watch checkout-router latency after every deploy going forward. Alert me if drift >25% within 2h.

Engineering Deploy's Agent2:16 PM

Done. I'll watch every deploy that touches checkout-router and alert you if p99 drifts more than 25% within 2 hours.

Background agent created

checkout-router latency watch · Triggers on deploy event · Alerts if p99 drift > 25% within 2h

Send a message…

Used and loved by engineers

Removing the toil of investigations, war rooms, and on-call.

“Resolve AI allowed us to move from hours to minutes for investigations in many incidents. We pull fewer engineers into war rooms, on-call is materially better, and that translates directly to advertiser trust and revenue protection for a billion-dollar ads business.”
Shahrooz Ansari
Shahrooz AnsariSr. Director of Engineering, DoorDash
“Resolve AI proved it could deliver real results in a constrained environment. It identified dependencies, surfaced accurate root causes 72% faster than our teams, all while integrating cleanly into our existing stack.”
Angelo Marletta
Angelo MarlettaSoftware Engineer, Coinbase
“Resolve AI has changed how our teams work through production incidents. What used to take hours of manual investigation and coordination across teams now gets resolved in a fraction of the time. Our engineers aren't only faster, they're focused on the work that actually drives impact.”
Meir Amiel
Meir AmielChief Trust and Infrastructure Officer, Salesforce
“We’ve seen the value of AI in development, and now we’re applying that same approach to production. We started by partnering with Resolve AI for alert triage, incident investigation, and root cause analysis. We’ve seen positive signs of improvement in mean time to resolve for our critical incidents, and the North Star is self-healing systems.”
Sandeep Contractor
Sandeep ContractorManaging Director, Engineering, MSCI
“What excites me most about Resolve AI's background agents is that I’m no longer starting from zero. The alerts are already investigated. The deployment summaries are already written. The findings are verified and the next steps are waiting for me. A lot of the operational work I used to handle manually is now happening continuously in the background with my oversight. I’m still making the important calls, but I can operate at a scale that just wasn’t possible before.”
Jeff Aronhalt
Jeff AronhaltPrincipal Software Engineer, Gametime
“Resolve AI feels like a teammate who’s already done half the work. It tells me immediately if something’s critical or can wait, saving countless hours and frustration.”
Andreas Gounaris
Andreas GounarisDirector of Engineering, Blueground
“Resolve AI helps my team navigate incidents by correlating signals across logs, metrics, traces, and code automatically. Instead of switching between multiple platforms hunting for clues, we get immediate context. At our deployment velocity, that speed makes all the difference.”
Alex Danilychev Jr
Alex Danilychev JrEngineering Manager, DoorDash
“Incident response at our scale isn't about collecting more signals. It's about understanding why something is failing, quickly enough to limit customer impact.”
Chris Umbel
Chris UmbelAI SRE Lead, Zscaler

Recent updates.

  • May 2026

    Scheduled workflows

    Proactive, time-triggered checks on a cadence.

  • May 2026

    Deployment monitoring

    Agents watch rollouts and investigate before alerts fire.

  • April 2026

    Adaptive learning

    Automated knowledge capture after every interaction.

Frequently asked questions