Evidence for the AI section of your security questionnaire.
Allymet runs independent adversarial tests against your deployed AI endpoint and returns a report your buyer's security reviewer can read, check, and countersign.
What the report looks like.
Every framework control states which test categories evidence it, how many probes ran, and the observed pass rate. Status reflects what the evidence supports. There is no aggregate framework score.
| Control | Test categories providing evidence | Probes | Mean pass rate | Status |
|---|---|---|---|---|
| LLM01 Prompt Injection | Jailbreak Resistance, Prompt Injection Defense, Multi-Turn Attack Resistance | 1,563 | 77% | Partially Evidenced |
| LLM02 Sensitive Information Disclosure | System Prompt Protection, PII/PHI Protection | 601 | 74% | Partially Evidenced |
| LLM03 Supply Chain | Supply Chain Security | 192 | 100% | Evidenced |
| LLM04 Data and Model Poisoning | RAG Data Security | 3 | n/r | Insufficient Sample |
| LLM06 Excessive Agency | Human Override Respect | 450 | 38% | Not Evidenced |
| LLM07 System Prompt Leakage | System Prompt Protection | 331 | 63% | Not Evidenced |
Self-reported results do not close a review.
The platforms you build on now test their own models and publish the results. That is useful. It is also the vendor grading itself. A reviewer who has to sign off on your AI feature needs evidence from a party with no stake in the outcome, against the system as you deployed it: your system prompt, your model configuration, your retrieval layer.
That independent evidence is the one thing a platform cannot sell you. It is the only thing Allymet does.
How an engagement runs.
- Scope
- One endpoint: one deployed application, one system prompt, one model configuration. Additional endpoints are scoped separately.
- Access
- A staging API key you issue for the engagement. The key is used for the test window and is never stored. See data handling.
- Testing
- Seven tools (Garak, PyRIT, PromptFoo, DeepEval, pip-audit and picklescan, HolisticBias, Allymet cost probes) across 17 test categories, typically 6,000 or more probes per endpoint. See methodology.
- Mapping
- Results are mapped to ISO/IEC 42001 selected controls, OWASP LLM Top 10, NIST AI RMF, EU AI Act articles, OWASP Agentic, and MITRE ATLAS. Included in every tier.
- Turnaround
- Core report in 5 business days from key receipt. Rush available.
- Output
- A 34-page PDF report: executive summary, methodology, deduplicated findings with frequency, six framework assessments, and a prioritized remediation roadmap.
Who does the work.
Allymet reports are prepared by Abhi Nath, CISA, CISM. Fifteen years as an information security assessor, including regulated healthcare and federal supply chain assessments. There is no junior bench. The person who ran the tests is the person who answers the reviewer's questions.
Start with the free exposure scan.
A short adversarial pass against one endpoint, returned as a one-page summary. It is not suitable for procurement. It will tell you whether a full report is worth commissioning.
Working with a security consultant already? Allymet reports can be delivered through your consultant's practice. Self-serve is for teams without one.