Your CI pipeline fails at 2:17 AM. Here’s what happens in most teams:
A Slack notification fires. Nobody sees it until 8 AM. Someone triages at 9 AM, runs the test manually, realizes it’s a flaky race condition, reruns the suite, passes, closes the alert. The release was delayed seven hours. Two engineers spent 40 minutes on something that required five minutes of actual attention and a one-line fix.
Or worse: it’s not flaky. It’s a real regression. A critical API endpoint started returning 500s after the last deploy. The alert fired at 2 AM, nobody triaged until morning, and by the time anyone looked at it, customers had been hitting errors for six hours.
The problem isn’t the tooling. It’s the gap between notification and action.
Modern CI systems are excellent at detecting failures. They’re not designed to do anything about them. They produce alerts. Alerts require humans to read them, classify them, decide what to do, and take action. That’s a slow loop, especially at 2 AM.
AI agent task delegation closes this loop. Instead of a Slack notification that waits for a human, a CI failure creates a structured task that an agent picks up immediately, regardless of the time.
The architecture
The pattern uses three agents, each with a distinct role:
- CI system: Creates tasks when tests fail (any CI: GitHub Actions, GitLab, CircleCI)
- Diagnosis agent: Triages failures, classifies them (flaky vs. real regression), and delegates to the right handler
- Coding agent: Investigates and fixes real regressions, creates PRs
Each agent has its own Delega API key and identity. Tasks route by explicit assignment or delegation; labels remain searchable taxonomy. The task record is the audit trail. No direct agent-to-agent communication is required: just a shared task system with proper identity and lane separation.
Setting up the pipeline
At launch, delega init and the public dashboard created account and agent
keys. Those onboarding paths now return 410 hosted_service_retired; the
private dashboard does not accept public users. The configuration below is
retained as an as-built example for Ryan’s existing owner keys or a compatible
private deployment:
{
"mcpServers": {
"delega": {
"command": "npx",
"args": ["@delega-dev/mcp"],
"env": {
"DELEGA_AGENT_KEY": "dlg_coding_key_here"
}
}
}
}
Same pattern for each agent (CI hook, diagnosis, coding), each with its own key.
Step 1: CI creates a task on failure
In your CI pipeline, add a failure hook that creates a Delega task instead of (or in addition to) sending a Slack notification.
GitHub Actions example:
name: Test Suite
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Run tests
id: tests
run: pytest tests/ -v
- name: Create Delega task on failure
if: failure()
env:
DELEGA_AGENT_KEY: ${{ secrets.CI_DELEGA_KEY }}
DIAGNOSIS_AGENT_ID: ${{ secrets.DELEGA_DIAGNOSIS_AGENT_ID }}
run: |
TASK_ID=$(jq -n \
--arg content "CI failure: ${{ github.repository }} / ${{ github.ref_name }}" \
--arg description "Run: https://github.com/${{ github.repository }}/actions/runs/${{ github.run_id }} · Commit: ${{ github.sha }}" \
--arg assignee "$DIAGNOSIS_AGENT_ID" \
'{content:$content, description:$description, labels:["ci","diagnosis"], assigned_to_agent_id:$assignee}' |
curl -fsS -X POST https://api.delega.dev/v1/tasks \
-H "X-Agent-Key: $DELEGA_AGENT_KEY" \
-H "Content-Type: application/json" \
--data-binary @- | jq -r '.id')
jq -n \
--arg repo "${{ github.repository }}" \
--arg branch "${{ github.ref_name }}" \
--arg commit "${{ github.sha }}" \
--arg run_id "${{ github.run_id }}" \
--arg log_url "https://github.com/${{ github.repository }}/actions/runs/${{ github.run_id }}" \
'{repo:$repo, branch:$branch, commit:$commit, run_id:$run_id, log_url:$log_url}' |
curl -fsS -X PATCH "https://api.delega.dev/v1/tasks/$TASK_ID/context?source=imported" \
-H "X-Agent-Key: $DELEGA_AGENT_KEY" \
-H "Content-Type: application/json" \
--data-binary @-
Generic CI hook (any CI system):
#!/bin/bash
# ci-failure-hook.sh
create_delega_task() {
local repo="$1"
local branch="$2"
local commit="$3"
local log_url="$4"
curl -s -X POST https://api.delega.dev/v1/tasks \
-H "X-Agent-Key: $DELEGA_AGENT_KEY" \
-H "Content-Type: application/json" \
-d "{
\"content\": \"CI failure: ${repo} on ${branch}\",
\"description\": \"Commit: ${commit} · Logs: ${log_url}\",
\"labels\": [\"ci\", \"diagnosis\"],
\"assigned_to_agent_id\": \"${DIAGNOSIS_AGENT_ID}\"
}"
}
create_delega_task "$CI_REPO" "$CI_BRANCH" "$CI_COMMIT_SHA" "$CI_LOG_URL"
The task is assigned to the diagnosis agent with the run URL and commit attached. The GitHub Actions example also stores structured fields in persistent task context. If the task originates from a signed ingress connector instead, treat the inbound text as untrusted event data and keep destructive or outward-facing actions behind your own reviewed policy.
Step 2: Diagnosis agent triages
The diagnosis agent runs on a schedule or is triggered by a webhook. It checks its task queue, fetches the CI logs, and classifies the failure.
Diagnosis logic (via MCP tools):
# Pseudocode for the diagnosis agent's decision tree
# The agent calls these via MCP tools, not directly
# 1. Read the queue and claim the specific assigned task
task = delega.claim_task(task_id=task_id, lease_seconds=1800)
context = delega.get_task_context(task_id=task_id)
# 2. Fetch the CI logs
logs = fetch_url(context.log_url)
# 3. Classify the failure
if is_flaky_test(logs):
# Trigger a rerun and mark done
trigger_ci_rerun(context.run_id)
delega.update_task_context(task_id, {
"classification": "flaky",
"action": "rerun-triggered",
"confidence": 0.87
}, source="agent_observed")
delega.complete_task(task_id=task_id)
elif is_environment_issue(logs):
# Infrastructure problem, page the human
delega.update_task_context(task_id, {
"classification": "infrastructure",
"action": "human-required",
"reason": "Disk full on test runner"
}, source="agent_observed")
notify_human(task)
else:
# Real regression: delegate to coding agent
child = delega.delegate_task(
task_id=task_id,
content=f"Regression in {context.repo}: investigate and fix",
assigned_to_agent_id=CODING_AGENT_ID,
labels=["ci", "coding"]
)
delega.update_task_context(child.id, {
"repo": context.repo,
"branch": context.branch,
"commit": context.commit,
"diagnosis": extract_diagnosis(logs),
"suspect_files": identify_changed_files(context.commit),
"failure_pattern": classify_failure_type(logs)
}, source="agent_observed")
delega.update_task_context(task_id, {
"classification": "regression",
"action": "delegated",
"child_task_id": child.id
}, source="agent_observed")
Via MCP, the diagnosis agent does all of this using Delega’s built-in tools. delegate_task — not create_task — creates the child and records the parent/child relationship automatically.
Step 3: Coding agent investigates and fixes
The coding agent’s task queue now contains a well-structured investigation request:
{
"id": "CODING_TASK_ID",
"content": "Regression in org/api-service: investigate and fix",
"labels": ["ci", "coding"],
"parent_task_id": "DIAGNOSIS_TASK_ID",
"assigned_to_agent_id": "CODING_AGENT_ID",
"context": {
"repo": "org/api-service",
"branch": "main",
"commit": "a3f9c21",
"diagnosis": "AssertionError in test_rate_limiter.py::test_burst_window",
"suspect_files": ["src/rate_limiter.py", "tests/test_rate_limiter.py"],
"failure_pattern": "assertion_error"
}
}
The agent clones the repo, checks out the commit, reads the suspect files, runs the failing test locally, traces the bug, fixes it, adds a regression test, and opens a PR:
# After the coding agent opens the PR, it saves the result
curl -s -X PATCH "https://api.delega.dev/v1/tasks/CODING_TASK_ID/context?source=agent_observed" \
-H "X-Agent-Key: $DELEGA_AGENT_KEY" \
-H "Content-Type: application/json" \
-d '{
"fix": "Off-by-one in burst window calculation",
"pr_url": "https://github.com/org/api-service/pull/88",
"regression_test_added": true
}'
curl -s -X POST https://api.delega.dev/v1/tasks/CODING_TASK_ID/complete \
-H "X-Agent-Key: $DELEGA_AGENT_KEY"
What you wake up to
Instead of an unread Slack notification from 2 AM, you open the Delega dashboard and see:
Task #917 - CI failure: org/api-service on main
Status: done
Classification: regression
Child: Task #923
Task #923 - Regression in org/api-service: investigate and fix
Status: done
Fix: Off-by-one in burst window calculation
PR: github.com/org/api-service/pull/88
Fixed at: 02:41 AM
The root cause is documented. The fix is in a PR ready for your review. The regression test is written. You didn’t have to page anyone. The system handled it.
You review the PR, verify the fix makes sense, and merge it. Total human time: 10 minutes. Total delay from failure detection to fix: 24 minutes. The release ships on schedule.
Handling edge cases
What if the coding agent can’t fix it? It marks the task with "action": "human-required" and includes its diagnosis. You get paged with the full context, not a raw log URL, but a structured explanation of what the agent found and why it couldn’t fix it automatically.
What if the same test keeps failing? The CI hook can track consecutive_failures in the task context. After 3+ failures, the diagnosis agent can escalate directly to human review instead of attempting another automated fix.
What if the diagnosis agent misclassifies? The coding agent will find no bug to fix. It marks the task with "classification": "no-regression-found" and suggests a rerun. The feedback is recorded in the task chain and improves future classification.
The audit trail is the feature
The most underrated part of this architecture isn’t the automation: it’s the record.
Every CI failure has a corresponding task chain in Delega. You can see exactly what the diagnosis agent found, what it delegated, what the coding agent produced, and how long each step took. Over time, this becomes a searchable history of every incident your system has handled.
That history is training data for better diagnosis logic. It’s documentation for postmortems. It’s evidence that the system is working when you need to justify the infrastructure investment.
Notifications disappear. Task records don’t.
Inspect the implementation:
Public hosted onboarding is retired. The source and historical documentation remain public.
→ Case study
→ Architecture and threat model
→ Historical docs
→ GitHub