# Mail, spend and security detections

Source: https://codeherder.com/docs/self-host-detections/

Set SES mail alarms, budget and anomaly alarms for spend outside AI accounting, and map common security detections to event types for your SIEM.

This page covers three things your organization operates outside CodeHerder: mail delivery health, infrastructure spend, and security detections in your SIEM. The sections are independent. Read the one you need.

## Mail delivery (SES only)

This section applies only when your server sends mail through Amazon SES. If you use another mail service, watch that service for the same signals.

The server selects SES when `CH_MAILER=ses`. It also selects SES when `CH_MAILER` is unset and a Cognito user pool is set. Set these values:

- `CH_MAILER_FROM` is the sender address. It is required.
- `CH_SES_REGION` is the SES region. It defaults to `CH_AWS_REGION`, then to `us-west-2`.

The server sends the sign-up code, the email-change code, the invitation notice and the synthetic-check mail. Cognito sends its own sign-in mail. Set the user pool’s email configuration to SES (developer mode) with the same verified sender. Watch that path with the same alarms.

### Permissions

The server role needs `ses:SendEmail` on the verified identity. It needs no other SES action.

### Verify before launch

- The sender identity is verified. Use a domain identity with DKIM, SPF and DMARC.
- The identity lives in the same region as `CH_SES_REGION`.
- The account is out of the SES sandbox.

### Alarms to set

Set these alarms in CloudWatch. Route them to your instance operator. The values are starting points. Tune them to your volume.

| Signal | Source | Warn | Page |
| --- | --- | --- | --- |
| Bounce rate | `Reputation.BounceRate` in `AWS/SES` | 2 % | 5 % |
| Complaint rate | `Reputation.ComplaintRate` in `AWS/SES` | 0.05 % | 0.1 % |
| Send failures | A CloudWatch Logs metric filter on the server log for `ses send email` | Not used | Any in 15 minutes |
| Suppression-list additions | An SES event destination for Bounce and Complaint, sent to SNS or EventBridge | Each addition | Not used |

### When sign-in mail fails

1. Run `ch instance check`. The `mail` check tells you if the server can send. See [Self-hosted synthetic checks](https://codeherder.com/docs/self-host-checks/).
2. Check the suppression list for the address: `aws sesv2 get-suppressed-destination --email-address <address>`. Remove the address only after you fix the cause.
3. Check the sending status and quotas: `aws sesv2 get-account`.
4. Check the Cognito user pool’s email configuration.
5. As a stop-gap, mint an API key for the person. See [Credentials and profiles](https://codeherder.com/docs/credentials/).

The send path does not retry. The person must ask for a new code.

## Spend outside AI accounting

CodeHerder cost accounting records AI usage only. See [Costs](https://codeherder.com/docs/costs/). It does not record these costs:

- The app host: EC2 and EBS.
- RDS PostgreSQL, its backups and its snapshots.
- The S3 attachments bucket.
- Data transfer and NAT gateway charges.
- CloudWatch logs and metrics.
- KMS keys and requests.
- SES and Cognito.
- EC2 device hosts and any microVM runners.
- The AWS bill for Bedrock. CodeHerder estimates Bedrock tokens from its price table. The AWS bill is the authority.

Set two alarms in the AWS account that holds these resources:

1. **An AWS Budget.** Use one monthly cost budget. Scope it by a cost-allocation tag, or use a dedicated account. Alert on forecast at 80 % and on actual at 100 %.
2. **A Cost Anomaly Detection monitor.** Monitor AWS services, or the same tag. Add an alert subscription for an impact above a dollar amount you choose.

Send both alerts through SNS to your finance operations team. Copy your instance operator. A runaway launch of device hosts or runners is the fastest way to a large bill, so review these alarms after any change to runner limits.

### AI bills CodeHerder only estimates

CodeHerder records an estimate for some AI spend, and nothing for the rest. Check each item against the provider’s own bill.

- **Anthropic direct.** CodeHerder computes `usd_micros` from its price table, not from your invoice. An optional daily reconcile compares the two. It needs `ANTHROPIC_ADMIN_API_KEY` and `ANTHROPIC_ORG_ID`. See [Costs](https://codeherder.com/docs/costs/).
- **Your gateway.** The gateway bills you on its own invoice. CodeHerder sees tokens, not the gateway’s fees or markup.
- **Subscription seats.** A seat is a flat fee. Its rows carry `billing_class = subscription`, and the dollar figure is notional.
- **Unpriced models.** A model with no price records $0 under the `__unpriced__` sentinel. Open `/instance/cost-pricing` to see the unpriced report. A $0 row is not free usage.

### Example alarms

This Terraform example creates both alarms for one deployment. It also adds a second budget for Bedrock, because the Bedrock bill lands in your AWS account. Change the limits, the tag and the address. Budgets and Cost Explorer are global services, so use the `us-east-1` provider region.

```
terraform {
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.0"
    }
  }
}

# Budgets and Cost Explorer are global services. Use us-east-1.
provider "aws" {
  region = "us-east-1"
}

variable "finops_email" {
  type        = string
  description = "Address that receives spend alerts."
}

resource "aws_sns_topic" "spend" {
  name = "codeherder-spend-alerts"
}

data "aws_iam_policy_document" "spend_publish" {
  statement {
    actions   = ["sns:Publish"]
    resources = [aws_sns_topic.spend.arn]

    principals {
      type        = "Service"
      identifiers = ["budgets.amazonaws.com", "costalerts.amazonaws.com"]
    }
  }
}

resource "aws_sns_topic_policy" "spend" {
  arn    = aws_sns_topic.spend.arn
  policy = data.aws_iam_policy_document.spend_publish.json
}

resource "aws_sns_topic_subscription" "finops" {
  topic_arn = aws_sns_topic.spend.arn
  protocol  = "email"
  endpoint  = var.finops_email
}

# Whole CodeHerder deployment, selected by a cost-allocation tag.
resource "aws_budgets_budget" "codeherder" {
  name         = "codeherder-monthly"
  budget_type  = "COST"
  time_unit    = "MONTHLY"
  limit_amount = "5000"
  limit_unit   = "USD"

  cost_filter {
    name   = "TagKeyValue"
    values = ["user:app$codeherder"]
  }

  notification {
    comparison_operator       = "GREATER_THAN"
    threshold                 = 80
    threshold_type            = "PERCENTAGE"
    notification_type         = "FORECASTED"
    subscriber_sns_topic_arns = [aws_sns_topic.spend.arn]
  }

  notification {
    comparison_operator       = "GREATER_THAN"
    threshold                 = 100
    threshold_type            = "PERCENTAGE"
    notification_type         = "ACTUAL"
    subscriber_sns_topic_arns = [aws_sns_topic.spend.arn]
  }
}

# Bedrock tokens bill to this account, so watch that service alone.
resource "aws_budgets_budget" "bedrock" {
  name         = "codeherder-bedrock-monthly"
  budget_type  = "COST"
  time_unit    = "MONTHLY"
  limit_amount = "2000"
  limit_unit   = "USD"

  cost_filter {
    name   = "Service"
    values = ["Amazon Bedrock"]
  }

  notification {
    comparison_operator       = "GREATER_THAN"
    threshold                 = 80
    threshold_type            = "PERCENTAGE"
    notification_type         = "FORECASTED"
    subscriber_sns_topic_arns = [aws_sns_topic.spend.arn]
  }

  notification {
    comparison_operator       = "GREATER_THAN"
    threshold                 = 100
    threshold_type            = "PERCENTAGE"
    notification_type         = "ACTUAL"
    subscriber_sns_topic_arns = [aws_sns_topic.spend.arn]
  }
}

resource "aws_ce_anomaly_monitor" "services" {
  name              = "codeherder-services"
  monitor_type      = "DIMENSIONAL"
  monitor_dimension = "SERVICE"
}

resource "aws_ce_anomaly_subscription" "immediate" {
  name             = "codeherder-anomalies"
  frequency        = "IMMEDIATE"
  monitor_arn_list = [aws_ce_anomaly_monitor.services.arn]

  subscriber {
    type    = "SNS"
    address = aws_sns_topic.spend.arn
  }

  threshold_expression {
    dimension {
      key           = "ANOMALY_TOTAL_IMPACT_ABSOLUTE"
      match_options = ["GREATER_THAN_OR_EQUAL"]
      values        = ["100"]
    }
  }

  depends_on = [aws_sns_topic_policy.spend]
}
```

The example was checked with `terraform fmt -check` and `terraform validate`. Confirm the email subscription in the inbox, then send a test message to the topic.

Anthropic direct and a gateway bill outside AWS. AWS alarms cannot see them. Set the spend limit in the Anthropic console, and set the budget in the gateway’s own console.

### When a spend alert fires

Start with the smallest action. Move down the list when spend continues.

1. **Pause new spawns.** Run `ch instance spawn-pause pause --fleet --reason "spend alert"`. Use `--root <workspaceRef>` to pause one root. See [Pause new work](https://codeherder.com/docs/self-host-incident-containment/#step-1-pause-new-work).
2. **Freeze running sessions.** Run `ch instance spawn-pause freeze --fleet --reason "spend alert" --yes`. Allow 30 seconds. See [Freeze a root](https://codeherder.com/docs/self-host-incident-containment/#step-2-freeze-a-root).
3. **Revoke the provider keys at the provider.** This works even when you do not trust the server. Revoke the key in the Anthropic console. For Bedrock, deactivate the IAM access key, or remove `bedrock:InvokeModel` from the device role. Revoke the gateway key in the gateway. Then clean up in CodeHerder. Run `ch device credential disable <label>` or `ch device credential delete <label>` on the device. Run `ch variable unset <KEY>` for a key held as a secret variable.
4. **Stop runners and device hosts.** Run `ch workspace edit <workspaceRef> --runner-scaling false` for each workspace that launches runners. Deactivate the runner IAM key. Scale the device Auto Scaling group to zero.
5. **Suspect the server itself.** Follow [Replacing a compromised self-hosted server](https://codeherder.com/docs/self-host-compromise/).
6. **Undo.** See [Undo](https://codeherder.com/docs/self-host-incident-containment/#step-6-undo). Restore provider keys before you unpause.

## Security detections for your SIEM

CodeHerder gives your SIEM two kinds of source:

- **Webhook events.** Create a subscription with `--format ocsf` and `--payload-mode metadata_only`. See [Webhooks](https://codeherder.com/docs/webhooks/). Only subscribable event types reach a webhook.
- **Other records.** The instance audit table `admin_audit_log` (see [Self-hosted log retention and access](https://codeherder.com/docs/self-host-log-retention/)), the activity feed, and the reverse-proxy access log (see [Self-hosted logs](https://codeherder.com/docs/self-host-logs/)).

### Detection matrix

| Detection | Webhook event types | Other records | Receiver test |
| --- | --- | --- | --- |
| Credential misuse | `api_key.minted`, `api_key.revoked`, `device.token_revoked`, `secret.rotated` | Reverse-proxy 401 and 403 bursts per client. Activity feed only: `sso.scim_token_minted`, `sso.scim_token_revoked`. Expiry: `sso.credential_expiring`, `sso.credential_expired`. | Mint an API key with `ch`, then revoke it. Expect one alert for each event. |
| Privilege change | `workspace.member_role_changed`, `team.member_role_changed`, `team.member_added`, `invitation.accepted`, `agent.created` | Activity feed only: `device.operator_capabilities_changed`, `sso.connection_updated`. `admin_audit_log` rows for `/v1/instance/` writes. | Change a test member’s role, then change it back. Expect an alert with the member as subject. |
| Mass export | None | `admin_audit_log` rows for `/v1/instance/` reads. Reverse-proxy counts of GET list and attachment routes per credential. Activity feed only: `self_host.download_completed`. | Read the instance humans list with an operator key. Expect a rule on the `admin_audit_log` row. Run 200 list requests with one key. Expect the proxy-count rule. |
| Webhook tampering | `webhook.disabled`, `inbound_webhook.created`, `inbound_webhook.deleted`, `inbound_webhook.secret_rotated`, `inbound_webhook.rule_changed` | Outbound webhook create, update, delete, and a manual disable or enable emit no event. `webhook.disabled` fires only when delivery fails repeatedly, so a receiver made to fail can hide tampering. Diff `GET /v1/workspaces/{id}/webhooks` each day against a known list. | Create an inbound webhook, rotate its secret, then delete it. Expect three alerts. Point a test outbound webhook at an endpoint that returns 500. Wait for the auto-disable. Expect one `webhook.disabled` alert. |
| Device enrolment | `device.token_minted`, `device.workspace_linked`, `device.enabled`, `device.restored` | Activity feed only: `device.tunnel_connected`. | Mint a device token in a test workspace and link a device. Expect an alert for each event. |
| Erasure | None | The `admin_audit_log` row for `POST /v1/instance/humans/{id}/erase`. | Erase a test person. Expect a rule on the `admin_audit_log` row. |

A test in `support/self_host_detections_test.go` checks that every event type in the second column is a live, subscribable type in `internal/eventtype`.

### Test the receiver

1. Create the webhook subscription for the event types in the matrix.
2. Run `ch webhook test <id>`. Confirm the SIEM receives the test event.
3. Run the action in each row’s last column, in a test workspace.
4. Confirm the matching rule fires.
5. Reconcile against the activity feed with `GET /v1/workspaces/{id}/activity?cursor=`. Every event in the feed that the matrix lists must also exist in the SIEM. A missing event means a lost delivery. See [Webhooks](https://codeherder.com/docs/webhooks/) for retries and failed deliveries.

### Gaps

CodeHerder does not emit these today:

- Sign-in outcomes, successful or failed.
- Outbound webhook create, update and delete, and a manual disable or enable. The daily diff covers them.
- Bulk reads as one event.
- Erasure as an event. It appears only in `admin_audit_log`.
