Profile
Back to NewsBack
Dev.to 24 min
Reader Mode
theAuth Audit Trails: Compliance Logging for AI Agent Actions

theAuth Audit Trails: Compliance Logging for AI Agent Actions

12 hours ago

theAuth is open-source auth for AI agents and humans. Star the repo on GitHub · Read the docs · Run the quickstart · theauth.dev

Picture the question that lands in your inbox six months after launch. "Which of your agents touched customer record 4417 on March 3rd, and who approved it?" If your answer starts with "let me grep the app logs," you already know how the next two days go.

I have been on the receiving end of that kind of question. Agents make it worse than normal services, because an agent decides at runtime what to call. You cannot read the code and know what it did. You need a record written at the moment of each decision.

This guide builds that record from zero. We will use theAuth, an open-source TypeScript library for agent identity and authorization. By the end you will have a log that captures every agent decision, alerts that fire on suspicious behavior, a retention job, a GDPR path for user deletion, and an export you can hand to a reviewer.

I will also tell you what the library does not do. Some of it surprised me when I read the docs closely, and you should hear it before an auditor does.

This is guide 8 of 8 in the theAuth guides. It stands on its own, so you can start right here. It stands alone and pairs well with guides 6 and 7.

Guide Title Read it when
1 Add login to an existing Next.js app You have an app with no auth yet
2 Passwordless login: passkeys, links, OTP You want to drop passwords or add 2FA
3 Multi-tenant SaaS auth: orgs, RBAC, SSO, SCIM You sell to teams and companies
4 Migrate from Auth0 or Clerk You already run another provider
5 Give every AI agent its own identity You run AI agents and need to start somewhere
6 Cap agent spend and require human approval Your agents spend money or act on risky things
7 Secure an MCP server for production You expose tools over MCP
8 Build an audit trail for AI agent actions (this guide) Someone will ask what your agents did

Building for people? Start at guide 1. Building for AI agents? Start at guide 5. Every guide links to the docs page for each concept it touches.

TL;DR

Goal What you use Where it lives
Record every agent decision agents.auditAll (on by default) Audit trail
Find and filter decisions theauth.audit.query() Audit trail
Spot odd behavior Your own queries plus the onViolation hook Anomaly detection
Push alerts out Webhooks and the SSE event stream Webhooks
Expire old records theauth.audit.cleanup() Compliance
Handle deletion requests The gdpr plugin GDPR data rights
Hand evidence to a reviewer audit.export() or signed credentials Verifiable credentials

Expect roughly 45 minutes if you already have a Node project. The code runs on SQLite locally and Postgres in production.

Prerequisites

You need Node 20 or newer, a TypeScript project, and the package:

npm install @glinr/theauth

You also need one user ID that exists as a row in the users table. Every agent has an owner, and the owner column is a foreign key. The quickstart shows how to create that row if your auth provider lives elsewhere.

One honesty note before we start. theAuth gives you building blocks for compliance work. It does not make you compliant. The docs say so on three separate pages, and I agree with them. A control mapping is a reporting aid. Your processes, your legal review, and your counsel decide whether you meet an obligation.

Step 1: Create the instance and keep auditAll on

The audit log is a side effect of the permission engine. You do not write log lines by hand. You create the instance, and the engine writes a row each time you call authorize().

import { createTheAuth } from '@glinr/theauth';

export const theauth = await createTheAuth({
  database: { provider: 'sqlite', url: 'theauth.db' },
  agents: {
    enabled: true,
    maxPerUser: 10,
    auditAll: true, // the default, but I like seeing it written down
    tokenExpiry: '24h',
  },
});

The auditAll flag defaults to true. If you set it to false, the permission engine writes no rows at all, even for denials. That is a sharp edge, so I keep the flag explicit in the config where a reviewer can see it. The configuration reference lists every agent option.

Swap the database block for { provider: 'postgres', url: process.env.DATABASE_URL! } when you deploy. SQLite is fine on a laptop, and Postgres is the right choice once more than one process writes to the log.

Step 2: Give every agent an owner and a narrow grant

An audit entry is only useful if it names someone. Each agent has an ownerId, which becomes the userId on every row it writes. That link is the whole provenance chain: this action, by this agent, owned by this person.

const agent = await theauth.agent.create({
  ownerId: 'user-123',
  name: 'support-summarizer',
  type: 'autonomous',
  permissions: [
    { resource: 'mcp:tickets:*', actions: ['read'] },
    {
      resource: 'mcp:refunds:issue',
      actions: ['execute'],
      constraints: {
        requireApproval: true,
        maxCallsPerHour: 5,
      },
    },
  ],
});

// Shown exactly once. Put it in your secrets manager now.
console.log(agent.token);

Notice the shape of the grants. The agent can read tickets broadly, but the refund action is narrow, rate limited, and gated behind approval. Narrow grants make your log readable. If an agent can do everything, every row looks the same and the log tells you nothing.

The agent token appears once, at creation. theAuth stores only the SHA-256 hash. If you lose the token, you rotate it.

For the approval flow itself, see approval flows. One caveat from the quickstart: with requireApproval: true, authorize() denies until you wire up approval handling. The audit log records that denial like any other.

Step 3: Call authorize before every sensitive action

Now the part that fills the log. Call authorize() before the agent touches anything that matters.

const decision = await theauth.authorize(agent.id, {
  action: 'read',
  resource: 'mcp:tickets:4417',
  arguments: { fields: ['subject', 'status'] },
});

if (!decision.allowed) {
  throw new Error(`Denied: ${decision.reason} (audit ${decision.auditId})`);
}

The returned auditId equals the primary key of the row the engine wrote. Put it in your own application logs, in your trace spans, and in any error you surface. Later, a single ID joins your app logs to the audit trail.

If your service receives the raw bearer token from an incoming request, use authorizeByToken(token, { action, resource }) instead. It writes the same kind of entry.

What a row contains

Each entry carries the agent ID, user ID, action, resource, the parameters you passed, the result, a free-text reason on denials, the evaluation time in milliseconds, and a timestamp. The table also holds the IP address and user agent when the request carried them.

Two details matter for compliance work.

First, the parameters field stores the arguments you pass to authorize(). If you pass a customer's email address or a raw prompt there, it lands in your audit table. That is useful for investigations and a liability for privacy. I pass identifiers and field names, never free text from a user.

Second, the log does not record everything. The docs are explicit that calls rejected before the permission engine runs, such as an unknown or revoked agent, write no row. A revoked agent hammering your API leaves no trace in the audit table. Capture those at your gateway if you need them.

Step 4: Query the log

Everything you need for daily review comes from one method. Filters are optional and combinable, and results come back newest first.

const denials = await theauth.audit.query({
  agentId: agent.id,
  result: 'denied',
  since: new Date(Date.now() - 24 * 3_600_000),
  limit: 100,
});

for (const entry of denials) {
  console.log(entry.timestamp.toISOString(), entry.action, entry.resource, entry.reason);
}

Without a limit, the query returns every matching entry. On a busy table that is a bad day for your memory, so always set one in application code.

One gotcha worth a sticky note: the actions filter runs after limit and offset. A page can come back shorter than the limit you asked for. If you filter on action names, page through results until you get an empty page, and do not treat a short page as the end.

The same data is available over HTTP. The REST API reference documents GET /audit and the export endpoint. Read its warning twice: the adapters do not authenticate callers on most endpoints, including audit queries, so you must put your own authentication in front of the mount path. An open audit endpoint is its own incident.

If you would rather click than query, the admin dashboard has a filterable log viewer with export.

Step 5: Build anomaly detection from the log

Here is the part where I need to be direct. theAuth does not ship an anomaly detector, and the library has no scan() function. The anomaly page lists exactly what exists: a trust score with an anomalyCount factor, the onViolation hook, and a reserved anomaly.detected event name that nothing emits.

That is less than the word "anomaly" suggests, and I would rather you know now. The good news is that the audit log has everything a simple detector needs. Two queries cover most of what I care about.

Denial rate per agent

An agent with a high denial rate is probing, confused, or broken. Each case deserves a look.

export async function highDenialAgents(thresholdPct = 20) {
  const since = new Date(Date.now() - 24 * 3_600_000);
  const logs = await theauth.audit.query({ since, limit: 5000 });

  const byAgent: Record<string, { total: number; denied: number }> = {};
  for (const entry of logs) {
    const bucket = (byAgent[entry.agentId] ??= { total: 0, denied: 0 });
    bucket.total++;
    if (entry.result === 'denied') bucket.denied++;
  }

  return Object.entries(byAgent)
    .filter(([, b]) => b.total > 0 && (b.denied / b.total) * 100 > thresholdPct)
    .map(([agentId, b]) => ({
      agentId,
      denialRate: Math.round((b.denied / b.total) * 100),
    }));
}

I hard-code 20 percent as a default argument here. In your own code, read the threshold from an environment variable, because the right number depends on your traffic.

Call volume per hour

A runaway loop shows up as a spike in calls from one agent. Count the last hour and compare it with a ceiling you choose.

export async function callsLastHour(agentId: string): Promise<number> {
  const recent = await theauth.audit.query({
    agentId,
    since: new Date(Date.now() - 3_600_000),
  });
  return recent.length;
}

Run both from a scheduler every five or ten minutes. A cron job is enough. You do not need a streaming pipeline for this.

Why I do not trust anomalyCount

The trust score counts denied entries whose reason contains INSUFFICIENT_PERMISSIONS, privilege, or escalation. The built-in denial reasons are plain sentences like No permission grants agent "x" access to "write" on "y". None of those words appear, so anomalyCount stays at zero unless your own hook or wrapper writes a reason containing one of them.

Read the score with theauth.trust.computeScore(agent.id) if you want the denial rate factor and the last violation timestamp. Just do not present anomalyCount to an auditor as a privilege escalation detector. The trust scoring page explains how the levels work.

Step 6: Push alerts instead of polling

Queries tell you what happened. Alerts tell you while it happens. theAuth gives you three delivery paths, and they are separate mechanisms with separate event lists.

The onViolation hook

This hook fires for every denial returned by authorize(), with a coarse type such as permission_denied, rate_limited, ip_blocked, time_restricted, or approval_required. The library does not await hooks, and the docs warn that a throw inside one becomes an unhandled rejection. Wrap the body in try/catch.

import { createTheAuth } from '@glinr/theauth';
import { createEventStreamModule } from '@glinr/theauth/auth';

export const theauth = await createTheAuth({
  database: { provider: 'sqlite', url: 'theauth.db' },
  agents: { enabled: true, auditAll: true },
  hooks: {
    async onViolation({ type, agentId, action, resource, reason }) {
      try {
        stream.emit({
          id: crypto.randomUUID(),
          type: 'anomaly.detected',
          timestamp: new Date(),
          data: { violation: type, action, resource, reason },
          agentId,
        });
      } catch (error) {
        console.error('[audit] violation relay failed', error);
      }
    },
  },
});

const stream = createEventStreamModule({
  db: theauth.db,
  requireAuth: true,
  validateToken: async (token) => {
    const session = await theauth.auth.session?.validate(token);
    return session?.userId ?? null;
  },
});

The hook references stream before the line that creates it. That works because the hook only runs later, after both variables exist. Tidy it to taste.

Remember the warning on the event streaming page: the core never calls stream.emit() for you. Every event reaches the stream because your code emitted it. Here we emit anomaly.detected ourselves, which works even though the library never does.

The SSE event stream

The stream suits live dashboards and security tooling. Clients connect with a bearer token and receive named events. If a client drops, it can pass since to replay what it missed from the theauth_stream_events table.

const source = new EventSource(
  '/api/theauth/events/stream?token=your-token&types=anomaly.detected,agent.revoked',
);

source.addEventListener('anomaly.detected', (event) => {
  const payload = JSON.parse((event as MessageEvent).data);
  console.warn('Violation', payload.data);
});

Two sharp edges here, both from the docs. If you set requireAuth: true and skip validateToken, the module accepts any non-empty token. And automatic browser reconnect replays nothing when your event IDs are UUIDs, because the module parses the Last-Event-ID header as a date. Track the last timestamp yourself and pass since. The full details live on the event streaming page.

The replay table is separate from the audit table. Replay reads stream events, not audit entries, and returns at most 1000 per call, newest first. Do not mistake it for your system of record.

Webhooks for durable delivery

Use webhooks when an external system needs the event, such as a ticketing tool or a SIEM collector. Each endpoint has a URL, a secret, and an explicit event list.

const theauth = await createTheAuth({
  database: { provider: 'sqlite', url: 'theauth.db' },
  webhooks: [
    {
      url: process.env.OPS_WEBHOOK_URL!,
      secret: process.env.THEAUTH_WEBHOOK_SECRET!,
      events: ['agent.revoked', 'agent.rotated'],
      retries: 3,
    },
  ],
});

await theauth.agent.revoke(agent.id);
theauth.webhooks?.emit('agent.revoked', { agentId: agent.id });

The same rule applies: the core does not emit these for you. You call emit after the action. No wildcard exists, and the event list stays closed.

On the receiving side, verify the signature before trusting anything. The signature covers <timestamp>.<raw body>, not the body alone. Reject old timestamps to limit replays.

import { createHmac, timingSafeEqual } from 'node:crypto';

export function verifyWebhook(
  rawBody: Buffer,
  timestamp: string,
  signature: string,
  secret: string,
): boolean {
  const expected =
    'sha256=' +
    createHmac('sha256', secret).update(`${timestamp}.${rawBody.toString('utf8')}`).digest('hex');
  const a = Buffer.from(expected);
  const b = Buffer.from(signature);
  return a.length === b.length && timingSafeEqual(a, b);
}

Retries use exponential backoff of 1, 2, and 4 seconds. After the last retry, theAuth drops the delivery. It keeps no delivery record and offers no replay, and a process restart loses pending retries. For an audit pipeline, that matters. Treat webhooks as a notification channel, and treat the audit table as the record. If a webhook goes missing, the row still exists. The webhooks page has the full header and retry reference.

Which one when

Use onViolation to react inside your process. Use the stream for people watching a screen. Use webhooks to tell another system. None of them replaces a periodic query against the log, which is the only path with no delivery gap.

Step 7: Set retention on purpose

theAuth never deletes audit rows by itself. Nothing expires. The one deletion path is a method you call:

const { deleted } = await theauth.audit.cleanup({ retentionDays: 365 });
console.log(`Removed ${deleted} audit entries older than 365 days`);

That call permanently deletes rows older than the cutoff. It has no soft delete and no undo.

Pick the number from your obligations, not from disk anxiety. The compliance page notes that the EU AI Act sets its own log retention rules for high-risk systems, and tells you to confirm the period with counsel. I am not a lawyer, and I will not give you a number.

Three rules I follow.

Run cleanup from a scheduled job, never from request handling. Log the deleted count somewhere durable, so the deletion itself leaves evidence. And never run it during a live review or investigation. The compliance page has a warning about exactly that, and you only make that mistake once.

If your database supports partitioning or TTL, that is the better tool at scale. A single DELETE over millions of rows locks things you care about.

Step 8: Handle GDPR requests without breaking the log

Audit logs and privacy law pull in opposite directions. The log wants to remember. A user can ask you to forget. theAuth ships a gdpr plugin for this, and its options are worth understanding before you wire up a delete button.

import { createTheAuth } from '@glinr/theauth';
import { gdpr } from '@glinr/theauth/auth';

const theauth = await createTheAuth({
  database: { provider: 'postgres', url: process.env.DATABASE_URL! },
  agents: { enabled: true },
  plugins: [gdpr()],
});

The plugin registers three endpoints for the signed-in user.

Export

GET /auth/gdpr/export returns a JSON bundle with the user's profile, agents, sessions, audit events, delegations, organization memberships, and API keys. It gives you a summary, not a dump. It leaves out TOTP records, passkeys, OAuth tokens, approval requests, and budget policies, and each section holds summary fields only. Audit events and delegations appear only if the user owns at least one agent.

If your legal team reads "all personal data" literally, this endpoint alone will not meet that bar. You will need to extend it with your own queries.

Delete

DELETE /auth/gdpr/delete removes the account. It requires a literal confirmation string, which keeps accidental calls out.

const res = await fetch('/auth/gdpr/delete', {
  method: 'DELETE',
  credentials: 'include',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    confirm: 'delete my account',
    keepAuditLogs: true,
    deleteOrganizations: false,
  }),
});

The keepAuditLogs flag is the audit-versus-privacy dial. When true (the default), theAuth anonymizes audit rows rather than deleting them, so your history of decisions survives without pointing at a named person. When false, theAuth deletes the rows for that user.

Read the details before you rely on it. With keepAuditLogs: true, the agent rows stay with status revoked, and they still carry their name and metadata. If agent names contain personal data, anonymization is incomplete.

The Postgres warning

This one deserves its own heading. The docs state that the module toggles PRAGMA foreign_keys during deletion and audit anonymization. That statement is SQLite syntax. On Postgres or MySQL, the keepAuditLogs: true path and the final user deletion can fail with a database error unless your driver accepts the statement.

I cannot tell you from the docs whether your setup works. I can tell you to test deletion against your production engine, with a throwaway user, before a real request arrives. Do it this week.

Anonymize instead

POST /auth/gdpr/anonymize replaces the email with a deterministic anonymous value and clears the name, external ID, and metadata. It keeps the account row, agents, and audit history. It also removes TOTP records and passkeys. It does not revoke sessions, and it does not touch audit rows.

Choose it when you need referential integrity, for example to keep organization membership intact. Choose deletion when the user wants the account gone.

Finally, deletion is irreversible, and the plugin has no onBeforeDelete hook. Data you store outside theAuth that references the user ID becomes orphaned. Cascade those deletes in your own code first. The GDPR page also shows createGdprModule for background jobs and admin scripts.

Step 9: Export evidence for a reviewer

Eventually someone asks for proof. You have two export paths, and they answer different questions.

Plain JSON or CSV

import { writeFile } from 'node:fs/promises';

const q1 = await theauth.audit.export({
  format: 'json',
  since: new Date('2026-01-01'),
  until: new Date('2026-04-01'),
});

await writeFile('audit-2026-q1.json', q1, 'utf8');

The method returns a string, not a file, so you write it out yourself.

Remember the cap: exports return at most the 10,000 most recent matching entries. A busy quarter will exceed that without warning, and you would hand over a silently truncated record. Export in narrow windows and check the row count against the cap.

export async function exportByDay(start: Date, days: number): Promise<string[]> {
  const files: string[] = [];
  for (let i = 0; i < days; i++) {
    const since = new Date(start.getTime() + i * 86_400_000);
    const until = new Date(since.getTime() + 86_400_000);
    const raw = await theauth.audit.export({ format: 'json', since, until });
    const rows = JSON.parse(raw) as unknown[];
    if (rows.length >= 10_000) {
      throw new Error(`Window ${since.toISOString()} hit the export cap, split it further`);
    }
    files.push(raw);
  }
  return files;
}

Prefer JSON over CSV for anything a machine will parse. In the CSV writer, the writer wraps the reason column in quotes without escaping quotes inside it, and built-in denial reasons contain double quotes. A strict CSV parser can choke. JSON also includes the parameters field, which CSV drops. Neither format includes IP address or user agent, even though the table stores both.

Signed credentials

A JSON file is plain data that anyone could have edited. If a third party needs to check that nobody changed a record since you signed it, export the entries as W3C Verifiable Credentials.

import { exportAuditAsVC } from '@glinr/theauth/vc';
import { generateDidKey } from '@glinr/theauth';

const keyPair = await generateDidKey();
const records = await theauth.audit.query({
  since: new Date('2026-01-01'),
  until: new Date('2026-04-01'),
  limit: 5000,
});

const result = await exportAuditAsVC({
  since: new Date('2026-01-01'),
  until: new Date('2026-04-01'),
  issuerDid: keyPair.did,
  issuerConfig: {
    issuerDid: keyPair.did,
    privateKeyJwk: keyPair.privateKeyJwk,
    publicKeyJwk: keyPair.publicKeyJwk,
  },
  records,
  filter: (r) => r.result === 'denied',
});

console.log(`Signed ${result.count} credentials`);

In production, load a persistent issuer key instead of generating one on each run. Otherwise nobody can verify older exports against a stable DID.

Credentials expire after 24 hours, and the verifier rejects expired ones. Export close to the time you hand evidence over. And be clear about what a signature proves: it shows who signed a record and that it has not changed since. It does not show that the underlying log was complete. The verifiable credentials page covers the presentation output and the JWT format.

Step 10: Map it to the frameworks, carefully

The compliance page maps features to EU AI Act articles, NIST themes, SOC 2 criteria, and ISO 42001. It makes a good starting point for a conversation with your auditor. I would use it like this.

For record-keeping obligations such as EU AI Act Article 12, the audit table gives you automatic, per-decision records with identity, resource, action, parameters, and outcome. For human oversight such as Article 14, requireApproval, delegation depth limits, and revocation give you technical controls. For SOC 2 monitoring criteria, queries and exports give you evidence of review.

What the mapping cannot give you matters more. Tamper-evidence is the clearest gap. The SDK never updates rows, but a database administrator can. The docs recommend database-level audit features or shipping logs to an immutable store. If your auditor asks for tamper-evidence, ship a nightly export to object storage with versioning and retention locks, and keep the signed credential exports for the periods that matter.

Troubleshooting and gotchas

The log is empty after you ran a test. Check auditAll first. With the flag off, the engine writes nothing. Then check that the agent exists and is active, because unknown or revoked agents write no rows.

You get a UUID auditId but find no row. With auditAll: false, the engine still returns an ID without writing a row. Calls blocked by a beforeAuthorize hook, and calls for unknown agents, return an empty auditId.

Your action filter returns half a page. That is the post-pagination filter behavior described in Step 4. Page until empty.

Your export looks short. You hit the 10,000 row cap. Narrow the window.

Revoking an agent did not appear in the log. Revocation is an agent lifecycle change, not an authorization decision. Emit your own event, or record it in your own table, when you revoke.

You wanted to pause an agent and resume it later. Revocation is permanent. Create a new agent with the same configuration instead, and tag the replacement in your own records so the history stays connected.

Webhook signatures never match. You probably hash the parsed body instead of the raw body, or leaving out the timestamp prefix. Use the raw bytes, and join with a dot.

The stream replays nothing after a browser reconnect. Your event IDs are UUIDs. Pass since explicitly.

Deletion fails on Postgres. See the PRAGMA warning in Step 8, and test before you need it.

When this is not the right fit

Four honest limits, so you can decide quickly.

If you need cryptographic tamper-evidence on every row at write time, theAuth does not build that in. You would add a hash chain or an external immutable store.

If you need a ready-made anomaly detector with severity scores and a findings API, theAuth does not provide one. You build it from the queries above or bring a SIEM.

If you need guaranteed delivery of every event to an external system, webhooks drop after retries. Pull from the audit table on a schedule, and reconcile.

If you need a certification, you need an auditor, not a library.

Go deeper

The pages I kept open while writing this:

FAQ

Does theAuth log every action an agent takes?

No. It logs every authorization decision that reaches the permission engine, which means every authorize() and authorizeByToken() call against a known agent. If your agent never calls authorize() before acting, nothing gets recorded. The log also skips calls from unknown or revoked agents.

Can I make audit logs tamper-proof?

Not with the SDK alone. It never updates rows, but a database administrator still can. For tamper-evidence, use your database's audit features, ship logs to an immutable store, or export signed Verifiable Credentials for the periods you need to prove.

How long should I keep audit records?

That depends on the rules that apply to your system, and the docs tell you to confirm the period with counsel. The library keeps everything until you call audit.cleanup({ retentionDays }) or expire rows yourself in the database.

What happens to audit rows when a user asks you to delete their data?

With keepAuditLogs: true, the default, theAuth anonymizes the rows and keeps the agents with status revoked. With false, it deletes the user's rows. Test this on your production database engine first, because the module uses a SQLite-specific statement.

Does theAuth detect anomalies automatically?

No. The library ships no detector and no scan() function. You get the onViolation hook, a trust score, and the audit log. You write the detection queries yourself, and the examples above are a starting point.

Can I use this to prove SOC 2 or EU AI Act compliance?

It gives you evidence and a control mapping. It does not give you a certification or a compliance verdict. Treat the exports as inputs to your own process and your auditor's review.

Your turn

Which part of your agent stack would an auditor find hardest to explain today: who owns the agent, what it could do, or what it actually did? I would like to hear where the gap is for you.

Try it yourself

The fastest way in is the quickstart. If this guide saved you time, a star on GitHub helps other developers find the project, and the docs cover every option used above. More about the project lives at theauth.dev.

That is the last guide. Start again at guide 1, or browse the full docs. The full list sits in the table at the top of this page.


GDS K S · thegdsks.com · building Glincker · follow on X @thegdsks

An agent that cannot explain itself later is an agent you will eventually have to turn off.

Chat with me