Most of us learned attack paths on a network. Someone lands on a box, escalates
locally, then pivots sideways to the next box. Exploit, escalate, pivot, repeat.
It is a good model and it shaped a generation of tooling.
In a cloud account it is mostly the wrong model, and the tooling that inherited
it keeps looking in the wrong place.
The pivot in cloud is usually not an exploit. It is an API call that the
attacker is allowed to make.
A chain that uses exactly one vulnerability
Here is the shape of it, with the only memorable CVE at the very start:
- An app has an SSRF bug. Not even a good one. It fetches a URL you give it.
- You point it at the instance metadata service and read the credentials of the role attached to the instance.
- Those credentials let you call
sts:AssumeRoleon a role in the production account, because somebody wired up cross account access a year ago and the trust policy was never narrowed. - That production role can read the bucket holding database backups.
Count the exploits. One, at step one, and it is a web bug rather than anything
exotic. Steps two, three and four are the platform working exactly as
configured. There is nothing to patch at steps two, three or four, which is why
a vulnerability scanner has nothing to say about them.
Now sort your findings by CVSS and see where that SSRF lands. It is a medium.
Above it are a dozen criticals on hosts nothing can reach.
The edge that matters is the trust policy
An IAM role has two halves and people pay attention to the wrong one.
The permission policy says what the role can do. This is the half everyone
reviews, the half that linters check, the half that shows up in least privilege
talks.
The trust policy says who is allowed to become the role. This is the half
that creates the edge in the graph. A permission policy with s3:* on a role
nobody can assume is a non event. A tightly scoped permission policy on a role
that anything in the account can assume is a path.
So when you go looking, read trust policies first:
aws iam list-roles \
--query 'Roles[].{name:RoleName,trust:AssumeRolePolicyDocument}'
Things worth stopping on:
A principal that is a whole account. "AWS": "arn:aws:iam::111122223333:root"
does not mean the root user. It means every principal in that account that also
has permission to assume it. If 111122223333 is a sandbox where a lot of people
have sts:*, you have handed the sandbox a key to here.
A missing external ID on a third party role. If you gave a SaaS vendor a role
and the trust policy names their account without an sts:ExternalId condition,
then any other customer of that vendor can potentially ask the vendor's platform
to assume into your account. This is the confused deputy problem and it is still
everywhere.
Federated trust with loose conditions. OIDC and SAML trust is where this gets
genuinely hard to eyeball. A GitHub Actions OIDC trust with sub left as a
wildcard means any repository in your organisation, or in some misconfigurations
any repository anywhere, can mint credentials for that role.
Chains. Role A can assume role B, role B can assume role C. Each hop looks
reasonable in isolation. Nobody drew the whole thing.
This is a graph problem and it needs graph tools
You cannot find chains by reading policies one at a time. Three hops is already
beyond what anyone holds in their head, and real accounts have more.
What you want is every principal as a node, every trust relationship as an edge,
and then a walk from anything externally reachable to anything you would be sad
to lose. The complication is that the edges have preconditions: you only take the
IAM edge if you first got code execution on the thing holding the role, which is
why this is not a plain shortest path.
I ended up writing a tool for the trust half of it, because reading
AssumeRolePolicyDocument by hand across three providers got old. It is called
frontdoor, it is Apache 2.0, it maps
federated and cross account trust into AWS, GCP and Azure, and it is read only
and needs no account. One binary. Point it at your credentials and it prints who
can walk in from outside.
That covers the trust graph. Joining it to reachability and to your
vulnerability data is the harder half, and that is the product I work on rather
than the free tool, so take that part as disclosure rather than advice: build it
yourself if you have the time, because you will understand your own account
better than any tool will explain it to you.
The thing I would tell my earlier self
Stop asking "what is the highest severity finding". Start asking "what can an
attacker reach, and what can they do once they are there".
The second question has a much shorter answer, and it is the one you can actually
finish.
If you handle federated trust differently, particularly OIDC conditions across
multiple clouds, I would like to hear it. That is the part I am least confident
about.