Skip to content

VIII.5The Security Governancegovern

OWASP AIMA - running an AI maturity assessment

A maturity score is only worth having if two assessors would give the same organization the same number, and if the result tells someone what to fix first. This chapter is how to get both out of OWASP AIMA, the open model for measuring how well an organization builds, secures and runs AI.

VIII.4 · ISO/IEC 42001, verification & maturity places AIMA in the OWASP standards chain and offers a quick four-rung ladder for a first conversation. This chapter is the working method behind it: what the model contains, how to score it without fooling yourself, and what the output looks like in a client engagement and in an internal program.

What AIMA is, and what it is not

The OWASP AI Maturity Assessment (AIMA) is an open, community-built maturity model for AI. It adapts OWASP SAMM, OWASP’s long-established software assurance maturity model, to the AI lifecycle, and extends it to data provenance, model robustness, privacy, fairness and transparency. Version 1.0 was released on 11 August 2025 under project co-leads Matteo Meucci and Philippe Schrettenbrunner. It ships as a PDF and an Excel toolkit (v1.0.1), is licensed CC BY-SA 4.0, and is an OWASP Incubator project. Everything is in the project repository.

The project describes V1.0 as coming with “detailed criteria and a worksheet for internal use or third-party evaluations”. Those are the two uses this chapter covers. Two boundaries keep expectations honest:

  • It measures the program, not a system. AIMA asks whether the organization reliably does the things that keep AI safe: strategy, policy, data handling, threat modeling, testing, monitoring, incident response. Whether one model or agent is actually secure is a question for testing (VI.4 · AI red-team playbook) and for verification against requirements such as AISVS.
  • It is not a certification. Nobody certifies an organization “AIMA Level 2”. The certifiable standard is ISO/IEC 42001 (VIII.4 · ISO/IEC 42001, verification & maturity). AIMA tells you how far you are from being ready for that audit.

The model on one page

AIMA groups its practices into eight domains, three practices each. The right-hand column is the plain-language question an assessor is really asking.

DomainIts three practicesWhat you are really asking
Responsible AI PrinciplesEthical & Societal Impact · Transparency & Explainability · Fairness & BiasDo you assess who the AI affects, can you explain what it does, and do you find and fix bias?
GovernanceStrategy & Metrics · Policy & Compliance · Education & AwarenessIs there a strategy with owners and measures, a policy people follow, and training that reaches the people building AI?
Data ManagementData Quality & Integrity · Data Governance & Accountability · Data TrainingIs data quality controlled, is someone accountable for the data, and is training data collected, labeled and licensed properly?
PrivacyData Minimization & Purpose Limitation · Privacy by Design & Default · User Control & TransparencyDo you collect only what the use case needs, build privacy in by default, and give people control over their data?
DesignThreat Assessment · Security Architecture · Security RequirementsAre AI systems threat-modeled and built to written, verified security requirements?
ImplementationSecure Build · Secure Deployment · Defect ManagementAre build and release controlled for AI systems, and are flaws tracked to closure?
VerificationSecurity Testing · Requirement-based Testing · Architecture AssessmentDo you test AI systems, against their requirements and their architecture, on a schedule?
OperationsIncident Management · Event Management · Operational ManagementWould you detect an AI incident, respond to it with a plan, and run the system under control?

Each practice has three maturity levels and two streams. Stream A, which AIMA calls Create & Promote, is about doing the activity. Stream B, Measure & Improve, is about measuring it and acting on what you measure. In plain terms, Level 1 means the activity happens informally, Level 2 means it is defined, documented and repeatable, and Level 3 means it is embedded across the lifecycle and continuously improved.

One question sits at each level of each stream. This is the full grid for Threat Assessment, taken from the V1.0 toolkit:

LevelStream AStream B
1Is there basic awareness or informal identification of threats specific to AI systems?Are informal threat mitigation strategies occasionally discussed or implemented?
2Are threats systematically identified and documented for AI systems?Are documented mitigation strategies developed and periodically reviewed?
3Is comprehensive threat assessment consistently performed and integrated across AI lifecycle?Are proactive and comprehensive mitigation strategies continuously implemented and refined?

Six questions per practice, 24 practices, 144 questions. A lightweight pass is interviews and document review, not an audit.

Two scoring methods, and why you must pick one

The document and the toolkit score differently, and the difference is large enough to change a board slide.

The document method (gated, SAMM-style). Answer each question yes or no. A “yes” to every Level 1 question earns Level 1; a “yes” to every Level 1 and Level 2 question earns Level 2, and so on. Partial progress into the next level adds a ”+”, so a practice with every Level 1 criterion met and one of the two Level 2 criteria met scores 1+. Scores run 0, 1, 2, 3, with ”+” variants; 0 means no appreciable activity.

The toolkit method (averaged). The Excel toolkit asks for a number per question: 0 (“no maturity”), 1 (“initial maturity”), 2 (“not full maturity”) or 3 (“maturity”). The practice score is the sum of the six answers divided by six, so it comes out as a decimal. A Results sheet lists all 24 practice scores, and a separate sheet charts them. Each question also has an interview-note column for recording what you were told and what you saw.

Here is one practice, Strategy & Metrics, scored both ways from the same interview. The evidence: the strategy exists as an unapproved slide deck, the chatbot’s containment rate and escalation volume are reviewed monthly, nothing is measured on risk, and nothing links the AI strategy to the enterprise strategy.

QuestionDocument (yes/no)Toolkit (0-3)
Level 1, Stream Ayes3
Level 1, Stream Byes2
Level 2, Stream Ano1
Level 2, Stream Byes2
Level 3, Stream Ano0
Level 3, Stream Bno0
Practice score1+8 / 6 = 1.33

The two agree roughly here. They diverge when a team does advanced work without the basics: an organization with an automated Level 3 practice but no documented Level 2 process stays capped by the gated method, while the average pulls it up.

  • Use the gated method when the score will be compared year on year, reported to a board or regulator, or produced by a third party. One advanced team cannot inflate it.
  • Use the average for a fast internal baseline and for tracking a trend between formal assessments.
  • Either way, state the method in the report, and never compare a gated score with an averaged one.

Running an assessment, step by step

  1. Set the scope. The whole AI program, one business unit, or one AI product. When an activity is handled outside your scope, for example by a group-level privacy office, record that rather than scoring it “no”. AIMA’s own guidance warns against marking such items not applicable too early.
  2. Choose the depth. A lightweight assessment uses the worksheets, interviews and document review to give a provisional score. A detailed assessment adds verification: you sample the evidence to confirm the activity really happens, “not just paper compliance”. A practical mix is lightweight across all eight domains and detailed for the few practices leadership cares most about.
  3. Choose the scoring method and write it down, with how you read Level 1 and how you split overlapping evidence.
  4. Map questions to people. Each domain has natural owners, listed in the table below. Send the questions ahead so people bring artifacts, not opinions.
  5. Interview and collect evidence. Record what you were told and what you saw in the interview-note column. A policy that exists but nobody has read is a Level 1 answer, not Level 2.
  6. Score conservatively, then calibrate. No evidence means the lower level. Where two assessors are involved, run a calibration session on a few practices before finalizing, so “systematic” means the same thing across domains.
  7. Set targets from risk, not from the scores. Not every practice needs to reach 3. A company whose only AI is an internal writing assistant can accept Level 1 on Fairness & Bias; a lender using AI in credit decisions cannot.
  8. Turn the gaps into a roadmap. The gap between current score and target, per practice, is the backlog. Sequence it, give each item an owner and a quarter, and date the reassessment.
DomainWho usually answers
Responsible AI PrinciplesProduct owner of each AI system, legal or ethics lead, data science lead
GovernanceCISO or AI governance lead, risk and compliance, HR or learning for training
Data ManagementData owner or data governance lead, ML engineering lead
PrivacyData protection officer, product owner
DesignSecurity architect, ML or platform engineering lead
ImplementationPlatform or DevSecOps lead, ML engineering
VerificationApplication security or red team lead, QA
OperationsSOC and incident response lead, SRE or operations

Worked example 1: an assessment for a client

This scenario is illustrative.

A mid-size insurer runs two AI systems: a customer-support chatbot built on a hosted LLM with retrieval over policy documents, and an internal model that triages incoming claims. The CISO asks two questions: where do we stand, and what do we fix first?

Scoping note - agreed with the client before the first interview
Scope: the insurer's AI program - the support chatbot and the claims-triage model.
Outside scope (recorded, not scored "no"): the group privacy office; Privacy evidence comes from its reports.
Model: OWASP AIMA V1.0, all eight domains.
Depth: lightweight everywhere; detailed, with evidence sampling, for Governance, Verification and Operations.
Scoring: document method, gated yes/no with "+", because next year's result will be compared with this one.
Level 1 reading: at least an informal version of the activity exists.
Overlap: deployment evidence counts under Secure Deployment only.
Targets: set in a workshop after calibration, from the client's risk appetite.

Six of the 24 results, with the evidence behind each:

PracticeWhat the evidence showedScoreTargetFirst move
Strategy & MetricsUnapproved strategy deck; chatbot metrics reviewed monthly, nothing on risk1+2Approve the AI strategy with three risk indicators, reviewed quarterly by the risk committee
Policy & ComplianceAI acceptable-use policy approved and communicated; AI-related legal requirements documented and reviewed yearly22None this cycle; hold the level
Threat AssessmentChatbot threat-modeled at launch, claims model never; mitigations never reviewed12Threat-model every AI system at design and on material change (VI.3)
Security TestingYearly web pentest touched the chatbot’s interface and its findings were fixed; no prompt-injection or jailbreak testing12A documented AI test plan run every release (VI.4)
Incident ManagementIR plan covers IT outages only; one prompt-injection report handled in an email thread12An AI incident playbook plus one tabletop exercise (VII.3)
Fairness & BiasBias discussed informally; claims handlers spot-check a sample of triage decisions12Define fairness metrics for the claims model and test them before each retrain

What the client receives:

  • The scoping note above, so next year’s assessment can use the same rules.
  • The full profile, 24 current scores against 24 targets, one page.
  • The top three moves with owners and quarters. Here: the strategy with risk indicators, the AI test plan and the incident playbook, because each one closes a gap on both systems at once.
  • A reassessment date six to twelve months out, when the first moves should have landed.

Resist the temptation to recommend “Level 3 everywhere”. A roadmap that closes three gaps this year is worth more to the client than a slide showing 24 red cells.

Worked example 2: adopting AIMA inside your own organization

Also illustrative. A 300-person software company ships LLM features in its product and runs an internal coding assistant. The head of security wants one measure of the AI program that engineering will actually use.

  • Name an owner per domain, using the table above, and give each owner their six questions per practice.
  • Baseline with the toolkit average. It is fast, and the decimals make progress visible quarter to quarter. Keep the evidence for each answer in the interview-note column, so the next quarter’s answers can be checked against it.
  • Tie targets to work the teams already plan. Security Testing rises when AI tests run in CI (VII.1 · The secure AI SDLC); Event Management rises when model and tool calls reach the SIEM (VII.3 · Detection, IR & forensics for AI).
  • Reassess every quarter, and calibrate once a year. A yearly external or peer assessment, using the gated method, keeps the internal numbers honest.
  • Re-baseline when scope changes. Launching an agent product adds new surface; last quarter’s score for Security Testing may no longer describe the system you run.

What the trend looks like after three quarters (illustrative numbers, toolkit method):

PracticeQ1Q2Q3What changed
Security Testing0.831.502.00Injection and jailbreak probes added to CI, results tracked per release
Event Management0.500.831.67Model and tool-call logs shipped to the SIEM with two detections
Incident Management1.001.001.83AI playbook written and exercised once
Data Training1.171.331.33Dataset licensing review started, not yet systematic

Report the profile, not an overall average. A program average of 1.8 can hide a 0.5 in incident management, and that is the number that matters on the day something goes wrong.

Evidence that supports each level

AIMA defines the criteria, not the artifacts. This is typical evidence for four practices, as a starting point for your evidence request.

PracticeLevel 1Level 2Level 3
Strategy & MetricsA strategy note or deck, even unapproved; some AI metrics trackedAn approved strategy, communicated, with a defined set of risk indicators reviewed on a scheduleRisk indicators in board reporting; the strategy revised from what the metrics show
Threat AssessmentA threat list for at least one AI systemA documented threat model for every AI system, with mitigations reviewed periodicallyThreat modeling is a lifecycle gate, updated on every material change
Security TestingAn occasional assessment that touched an AI featureA documented AI test plan, covering injection, jailbreaks and data leakage, run on a scheduleAI tests run in CI; results tracked as metrics and audited
Incident ManagementAd hoc handling of an AI-related reportAn AI incident procedure used consistently, with incidents documented and reviewedProcedures exercised through tabletops; lessons fed back into design and testing

Where AIMA sits next to the other frameworks

AIMA V1.0 does not ship a formal mapping to other frameworks. The table below is this book’s reading of how they fit together.

If you also useWhat it answersHow AIMA fits
NIST AI RMF (VIII.3)The risk-management process: govern, map, measure, manageAIMA measures how consistently you perform the activities your RMF profile commits to
ISO/IEC 42001 (VIII.4)A certifiable AI management systemA readiness check before the audit, and the internal measure between audits; not a substitute for certification
EU AI Act (VIII.6)Legal obligations by risk tierProgram maturity, not legal compliance; map specific obligations separately
AISVS and AIVSS (VIII.4)Testable requirements, and the severity of a findingTest results become evidence for the Design, Implementation and Verification scores
The four-rung ladder (VIII.4)A quick headline for a first conversationMove to AIMA when the result must be repeatable and auditable