I build and deploy cloud infrastructure and AI agents on AWS.
Six systems, each one built and measured. Query them below, or open any one and read the architecture in full.
$ls systems/# click a row to open it
cloudfix
reviews a terraform plan before apply, returns a verdict with cited evidence
bedrock terraform python
cloudops-sentinel
stops idle ec2 on a schedule, least privilege enforced at two independent layers
eventbridge lambda terraform
support-agent
support agent with business rules in its tools and every tool call audited
bedrock strands agentcore
iac-pipeline
whole network and compute stack from one apply, zero long lived keys
terraform s3 backend github actions
k8s-observability
containerised agent on kubernetes, self hosted monitoring measured against datadog
docker kubernetes prometheus
diamond-velmora
paid client build delivered end to end, in production
s3 cloudfront route53
$cat next.txt
nextSupport agent on AgentCoreredeploy in my own account
nextAgents on Amazon EKSrebuild the workshop in my own account
plannedHashiCorp Terraform Associateexam
Not on the site as experience until it is.
$cat systems/cloudfix/README
CloudFix
terraform plan review, before apply · exits 0 safe / 1 review / 2 do-not-apply
A Terraform plan that destroys your database looks exactly like one that does not, until it has run. CloudFix reads the plan and returns SAFE, REQUIRES HUMAN REVIEW or DO NOT APPLY, with every claim cited back to the plan itself.
It exits 0, 1 or 2, so it drops into a pipeline as a gate rather than sitting somewhere as a report nobody opens.
DeterministicParsePlan JSON into resources and actions
DeterministicCheck7 functions, 16 finding types
DeterministicBlast radiusWhat dies and what depends on it
ModelAssessVerdict, every claim cited
DeterministicVerify5 rules. Do the citations resolve?
Model, on failureRepairFix the reasoning, verify again
Green stages are deterministic Python.Amber stages are the only two places the model is permitted to reason.
$cat decisions.md
The model never closes the loop
The obvious build lets the model read the plan and state a verdict. That fails quietly: a confident wrong answer looks exactly like a right one. Here the model's output is checked by deterministic code that resolves every citation back to the plan. If a citation does not resolve, the verdict is not trusted and a repair pass runs.
It holds no credential that can change anything
CloudFix has no write permissions at all. Not scoped write permissions, none. Whatever it concludes, it is structurally incapable of acting on it. That is a cheaper guarantee than trying to constrain behaviour through prompting.
Exit codes, not a report
Returning 0, 1 and 2 means CI can gate on it with no glue code. A tool that produces a nicely written document a human has to read has not actually removed the risk, it has moved it to whoever is too busy to read.
self healing infrastructure · two independent monitoring loops
An idle instance costs money all night. Sentinel runs two independent loops against live EC2. The heal loop fires every hour from an EventBridge schedule: a Lambda reads each running instance's average CPU for the past hour and stops any instance that is idle and tagged for auto stop. The alert loop routes a CloudWatch CPU alarm through SNS to a second Lambda that records it.
The whole heal loop is Terraform. The IAM role, the function, the rule, its target and the invoke permission all come up from nothing on one apply, with no configuration drift. The alert loop is partly there: its topic and role are imported into Terraform, and its function and alarm are next to be rebuilt as code.
Run the three scenarios below against the architecture. The third one is the interesting one.
EventBridgeScheduleFires every hour
AWS LambdaHealerChecks the tag itself
CloudWatchCPU metricsIdle below 5% over the hour
IAMPolicy conditionStop allowed only on the tag
Amazon EC2InstanceTarget of the stop call
The alert loop runs alongside: CloudWatch alarm, SNS topic, a Lambda that logs it.It records, it never stops anything.
/aws/lambda/auto_stop_idle_instancesidle
Pick a scenario. The output is simulated from the function's logic, not a live feed.
$cat decisions.md
Least privilege, twice, on purpose
An IAM policy condition scopes the permission to the instance tag, and the function independently checks the same tag before calling the API. That looks redundant until you consider that a bug lives in one layer at a time. Either check alone is a single point of failure. Run the third scenario and you can watch both refuse the same call.
Two loops rather than one clever one
Recording an alarm and stopping idle waste are different problems with different failure modes. Combining them into one function would have produced something harder to reason about and harder to test, for no gain.
The whole loop is code, including the permissions
It would have been faster to click the IAM role together in the console once. Then the next environment is a memory test. Everything, including the invoke permission, deploys from a single apply.
$cat results.json
monitoring_loops2
least_privilege_layers2
heal_loop_applies1
configuration_drift0
stack eventbridge lambda cloudwatch sns iam terraform
amazon bedrock, strands agents · deployed to agentcore in an aws workshop
A support agent on Amazon Bedrock, built with the Strands Agents SDK. It looks up orders, processes refunds and answers store policy questions through tool use, and the business rules live inside the tools rather than in the prompt.
During the AWS "Building with Amazon Bedrock" workshop I deployed it to AgentCore Runtime, with AgentCore Memory for cross session context and AgentCore Gateway exposing Lambda backed tools over IAM authenticated MCP. That ran in the workshop's temporary account, which no longer exists. The code in the repo runs locally with a Streamlit chat UI, and redeploying it in my own account is next.
StreamlitChat UIKeeps the conversation
Amazon BedrockStrands agentDecides which tool to call
Business rulesIn codeUnknown, refunded or unshipped orders refused
Event hookAudit trailEvery request and result recorded
Green stages are code. The model only chooses which tool to call. In the workshop deployment the same agent sat behind AgentCore Runtime, Memory and Gateway.
$cat decisions.md
Actions go through tools, not text
A refund is a function call with defined inputs and a fixed set of reasons. The model decides when to act, but not how the action is carried out.
Business rules live in the tool
The refund checks sit in code, so a persuasive customer cannot talk the model into refunding an order twice or refunding one that never shipped.
Every tool call is logged before anyone asks
An agent that can issue a refund needs an audit trail from day one, not after the first disputed transaction. Event hooks record each call, so any action can be traced afterwards.
A pipeline holding an AWS access key is a leaked AWS access key waiting for the repo to go public. This one authenticates through OIDC federation, so it receives a short lived token per run and no long lived credential exists anywhere to leak.
The full network and compute stack is code: VPC, public subnets across two Availability Zones, internet gateway, route tables, security group and EC2, reproducible from a single apply and split into reusable network and compute modules with separate dev, prod and global state.
Every pull request runs validate and plan. Apply runs only after merge, only when Terraform files change, and only after a person approves it.
GitHub ActionsPull requestvalidate and plan on every PR
The only non deterministic step in this pipeline is a human deciding to ship.
$cat decisions.md
Federation instead of stored keys
Storing an access key in CI secrets works immediately and is a liability forever. OIDC costs an afternoon of trust policy configuration and removes the entire class of problem.
Remote state with real locking
State on a laptop is fine until there are two laptops. The S3 backend is versioned, encrypted and locks natively, so two concurrent runs cannot corrupt each other.
A human gate in front of production
Everything up to production is automatic. Production waits for a person to press the button. Full automation into prod is a goal, not a starting position, and pretending otherwise is how you learn about rollbacks.
$cat results.json
long_lived_keys0
applies_for_full_stack1
required_checks_per_pr2
stack terraform s3-backend github-actions oidc iam vpc ec2
docker, kubernetes, prometheus, grafana · measured, then rebuilt from the runbook
An AI support agent that calls Amazon Bedrock, containerised with a multi stage Docker image that runs as a non root user, and run on a two node Kubernetes cluster. Prometheus finds it through pod annotations and scrapes its metrics, four Alertmanager rules watch it, and Grafana shows a dashboard provisioned from a JSON file in the repo.
Then the same workload under Datadog, to see what managed monitoring actually costs, and a full rebuild from a wiped machine using only the written runbook.
DockerImageMulti stage, non root
Kuberneteskind clusterTwo nodes, no cloud spend
The agent/metricsCounter and histogram, bounded labels
PrometheusScrape and alertFound by pod annotations, 4 rules
GrafanaDashboardProvisioned from a JSON file
Stated plainly: a local cluster, one replica of everything, metrics only.The load endpoint is synthetic, so its numbers are not application performance.
$cat decisions.md
The light chart, not the default stack
On a 4.8 GiB host, kube-prometheus-stack took 17 minutes and restarted pods 8 times. The light Prometheus chart installed in about 2 minutes with zero restarts. The right tool is the one that fits the machine it runs on.
Labels from route templates, not raw URLs
The endpoint label comes from the route template and falls back to "unmatched", so a scanner probing random URLs cannot mint a new time series for every URL it tries.
A runbook is real once it rebuilds from nothing
I wiped the machine and rebuilt the stack using only the README. It exposed five defects, including helm reporting "deployed" for a Grafana whose data source did not exist. All five are fixed in the runbook.
paid client engagement · delivered end to end, in production
A complete build for a Lagos spa and interior styling business: design, online booking with card deposit payments, a branded business email domain, AWS hosting and a launch campaign.
Paid client work is a different discipline to a personal project. The deadline is not yours to move and the consequences of getting it wrong belong to someone else.
Route 53Custom domainDNS to the distribution
ACMCertificateHTTPS enforced
CloudFrontGlobal CDNThe only route to the origin
S3 with OACPrivate originBucket never publicly reachable
Payments and mailThe business bitCard deposits, branded domain email
Same infrastructure discipline as everything above it. That is the argument for both halves being on one site.
$cat decisions.md
Origin Access Control, not a public bucket
The fast way to host a static site is to make the bucket public and point DNS at it. Then your origin is reachable directly, your CDN is decoration and your WAF rules protect nothing. OAC means CloudFront is the only path in.
Card deposits rather than free booking
A booking system with no deposit produces a calendar full of people who do not arrive. Taking a card deposit at booking was a business decision before it was a technical one, and it is why the system is still in use.
Branded email on a domain she controls
Running the business from a free mailbox means the business identity belongs to whoever owns that mailbox. Domain email with verified SPF and DKIM was part of the delivery, not an upsell.
I studied Biochemistry at the University of Lagos, then spent seven years leading operations for an events and lifestyle brand, managing delivery on more than fifty events and negotiating partnerships worth over ₦100M.
I moved into cloud engineering because I wanted to build systems rather than coordinate them. I earned the AWS Solutions Architect Associate certification in August 2026 and have been building on AWS since: first infrastructure, then agents, then containers.
The operations background is not incidental to the engineering. Knowing how delivery actually fails under a real deadline is why every system above has a second check behind the first one.