What is it?
Managing AWS environments means connecting technical work with decisions about service quality, security, cost and customer expectations. AWS infrastructure operations are part of that responsibility, alongside the application dependencies and people needed to deliver a usable service.
A new product needs capacity. A larger customer requires tighter access controls. Engineering introduces another service. These can all be signs of healthy growth.
The cloud you designed is not necessarily the cloud you are running. The problem begins when recovery instructions, ownership and operating procedures still describe an earlier version of the business.
Why it matters
Risks
- An undocumented dependency can invalidate a recovery procedure.
- In April 2022, an Atlassian maintenance operation accidentally deleted customer sites while removing a legacy app. The outage affected 775 customers, and some lost access for up to 14 days.
- Atlassian had backups and tested recovery, but lacked the capability to restore many interconnected customer sites at that scale. It met its recovery point objective but missed its recovery time objective.
Costs
- Capacity added for a busy period can keep running after demand falls.
Operational impact
- Alerts and workarounds accumulate, drawing engineers away from planned work.
Strategic impact
- Each new product becomes harder to support when responsibilities remain unresolved.
Five habits that keep AWS running well
Keep ownership current
- Give each important workload a named owner. Agree who investigates alerts, authorises changes and communicates with the business. Include a backup contact and review responsibilities when teams or services change.
- A payment alert might reach five people. Someone still needs authority to lead the response. Shared dashboards cannot resolve unclear decision rights, drawing on AWS Operational Excellence guidance that connects workload design, delivery and ongoing improvement.
Validate changes before and after release
- Keep infrastructure changes in version control, review them and test the behaviour they affect. A networking change may deploy successfully while breaking access to a dependency. Quality engineering connects the deployment result to the service the customer uses.
- AWS Config can record configuration history for supported resources, helping teams investigate what changed. That history complements reviewed infrastructure code and functional checks, but it does not establish that an application works correctly.
- Before deployment, agree how to recover from an unsuccessful change, including when to roll back or fix forward. Revisit the plan when dependencies or data changes make reversal more complicated.

Test recovery against today's dependencies
- A successful backup job confirms one part of the process. Recovery testing should establish whether restored data is usable and whether recovery meets agreed objectives.
- Take the exercise through the application: reconnect identity, networking and other dependencies, sign in, then complete a representative customer task. A database that starts successfully is a checkpoint. A usable service is the outcome.
- Measure recovery time against the recovery time objective (RTO), and data loss against the recovery point objective (RPO). Repeat relevant tests after significant changes, including growth that makes an earlier recovery scenario unrepresentative.

Watch customer experience and cost together
- Amazon CloudWatch Synthetics uses scheduled scripts called canaries to check endpoints and user actions. A sign-in check can expose a problem that infrastructure metrics alone do not explain. Keep those checks aligned with changing customer journeys.
- Alongside service health, review workload costs regularly. More customers may justify more capacity, while an abandoned test environment needs a different response. Where useful, track cost per transaction or active customer alongside performance.

Turn recurring problems into planned work
- AWS operations management needs time for improvement as well as response. Review operational measures, repeated incidents, manual work and permissions that no longer fit people’s roles.
- Give each priority an owner and a completion date. At the next review, check whether the change addressed the cause. As the business grows, yesterday’s acceptable workaround may deserve a permanent fix.
Common mistakes
Leaving launch documentation untouched
Consequence: New dependencies disappear from the operating picture.
Avoidance: Update runbooks and ownership as part of material changes.
Treating deployment success as service success
Consequence: Infrastructure can be healthy while a customer journey fails.
Avoidance: Validate the affected behaviour after release.
Reporting a problem without assigning the next action
Consequence: An alert, cost anomaly or failed test drifts without resolution.
Avoidance: Give each one someone responsible for resolving it.
Best practices
- Use a short service review to connect the five habits.
- Ask what changed, what customers experienced and which assumptions need retesting.
- Bring evidence appropriate to the workload: recent release checks, recovery results, cost trends and unresolved incidents.
- Keep the resulting action list small enough to complete.
- Let review frequency reflect service criticality and the pace of change.
How to get started
- Check the current dependencies, access requirements and escalation contacts.
- Identify one customer journey that must work and how you monitor it.
- Compare the latest release and recovery tests with today’s environment.
- Review cost against demand and identify unexplained changes.
- Assign the most consequential gap an owner and a completion date.
Choose one service the business depends on, and work with its business and technical owners. Start where failure would most affect customers or business operations. Resolve that gap before expanding the exercise across the platform.
How KineticSkunk helps
KineticSkunk’s Cloud Platform Engineers can build, modernise or take over an AWS platform, then operate an agreed scope through an ongoing managed relationship.
Our AWS managed cloud services connect platform operations with security, recovery, observability, cost optimisation and engineering assurance. Responsibilities, service boundaries and how we work with your team are agreed for the engagement. An AWS cloud management partner becomes useful when routine work consumes engineering capacity or important improvements keep slipping, so start by clarifying which responsibilities need consistent ownership.
Run AWS with the same discipline every quarter. Pair this with Five cloud mistakes that are holding fintechs back and Fintech cloud cost strategy when you want ownership, recovery and cost on the same review.
Frequently asked questions
AWS platform operations cover the ongoing work of running an AWS environment, including monitoring, incident response, access, changes, recovery and cost management.
New products, customers and dependencies change what the team must support. Monitoring, ownership, capacity and recovery procedures need to evolve with them.
Each important service needs accountable business and technical owners. Internal teams and partners can share tasks, with explicit decision rights and escalation paths.
No. Recovery also requires usable data, working application dependencies and tested procedures that meet the business’s downtime and data-loss objectives.
When operating responsibilities exceed available capacity, ownership is unclear or recurring issues displace improvement work. Agree the scope and responsibilities each team retains.
Sources
- Atlassian: April 2022 post-incident review
- AWS: Operational Excellence Pillar
- AWS: Resources have identified owners
- AWS: Plan for unsuccessful changes
- AWS: What is AWS Config?
- AWS: Perform periodic recovery of data
- AWS: CloudWatch Synthetics canaries
- AWS: Review and analyse workloads regularly
- AWS: Review operational measures and prioritise improvement
- AWS Managed Platform
- AWS Cost Optimisation & FinOps





