A symptom is not a root cause
Pending pods can result from taints, affinity, storage, insufficient regional capacity or an IAM denial. The quota diagnosis requires every modeled causal signal: Pending pods, failed NodeClaims, a recognized EC2 limit error and usage at or above the supported Standard On-Demand vCPU quota. Missing or contradictory signals produce an undetermined diagnosis, not a quota change.
The current input is a scoped synthetic snapshot with strict fields and a 15-minute freshness limit. It does not authenticate AWS or Kubernetes provenance. A live collector would need per-source timestamps, event identifiers, instance-family quota coverage and trusted account, region and cluster attribution before its output could support operational decisions.
Keep the action boundary deterministic
The proposal generator emits a fixed Terraform Service Quotas resource for the allowlisted Standard On-Demand EC2 quota. The target must exceed the current value and cannot exceed twice that value or 1024 vCPUs. It never runs Terraform, changes Kubernetes resources or creates a GitHub pull request.
An optional operator-supplied explanation JSON can describe the established evidence, but it cannot change the diagnosis or evidence IDs. This is an offline boundary, not a live model integration or a general prompt-injection defense. Accepted prose remains untrusted and does not select actions.
A request is not effective capacity
Submitting a quota request does not guarantee approval, and an approved quota does not guarantee an available EC2 instance. The demo remains awaiting recovery until a newer observation confirms the effective quota, successful launches, no Pending pods, no failed NodeClaims, cleared launch errors and application health.
SQLite transactions persist the investigation, proposal and verification history. Identical investigation and proposal replay is idempotent; conflicting evidence or targets are rejected. Audit events are appended by the application, but the database is local, unencrypted and writable by its owner—not an immutable compliance ledger.
Test the refusal paths, not just the happy path
- The 19-test suite covers malformed input, duplicate keys, oversized files, stale evidence and future timestamps.
- Alternative IAM and capacity failures must not produce quota proposals.
- Forged explanation evidence, changed incident scope and excessive quota targets are rejected.
- Each recovery signal is tested independently; a proposal alone cannot satisfy recovery.
- Ruff, Bandit and Python 3.12/3.13 CI checks complement the tests; they are not a penetration test or production certification.
Connect GitOps without making AI a dependency
The roadmap separates an Argo CD delivery plane from the investigation plane. Keycloak supplies operator identity, Harbor stores images, Vault supplies short-lived secrets, and admission policies would enforce approved registries and signatures. Those integrations are designed, not deployed by this release.
A future read-only investigator may propose a bounded PR through a separately scoped GitHub App. A human approves the complete plan; a protected workflow executes it; independent observations verify recovery over a stability window. Argo CD reconciles Kubernetes, while a separate Terraform workflow handles cloud quota requests. Manual alerts and the incident runbook must still function if every AI component is unavailable.
A useful SRE agent narrows uncertainty. It does not turn generated prose into authority or remove the operator's fallback path.