LLM Integration
Attest ID runs against a self-hosted, OpenAI-compatible inference endpoint (chat, vision, embeddings) to add AI-assisted features: issuer onboarding, bulk issuance, revocation, credential templates, catalog configuration, verifier UX, admin-console intelligence, and a docs/support assistant. None of these are load-bearing for correctness — every one degrades to “unavailable” cleanly when the model endpoint is unset, disabled, or unreachable.
Design decisions
- Core Engine only. Every LLM-assisted feature is a REST-facing workflow Core Engine already
owns.
LlmClientis wired once, inLlmConfig; Program Catalog, Credential Issuance, and Issuer Registry stay gRPC-only with no LLM dependency. - Never throws. Every
LlmClientmethod returns a typed outcome (COMPLETED/UNAVAILABLE/FAILED) instead of throwing — a call site checks availability and degrades (hide a UI section, fall back to manual entry). “Advise, never gate”: an LLM result informs a human decision and never auto-applies itself. - Structured output via JSON Schema, not prompt-engineered parsing — a call site gets a typed result back, not free text to regex out.
- Ephemeral documents. A feature reading an uploaded image (issuer KYC, holder ID cross-check) holds the bytes in memory for exactly one vision call, then discards them — only the structured extraction is persisted.
- In-memory retrieval, not a vector DB. Semantic catalog search and the docs/support assistant both embed a small, bounded corpus into an in-memory cosine-similarity index rather than taking a vector-DB dependency — revisit if either corpus grows past a few thousand items.
- Token budgets account for reasoning overhead. The configured chat model spends tokens on a hidden “thinking” phase before its answer — a budget sized only for the visible answer truncates the response before the model reaches it. Every feature budgets generously (400–600 tokens) even for short answers.
- Config:
attestpro.llm.*, all overridable viaATTESTPRO_LLM_*env vars.enabled: falseunregisters theLlmClientbean entirely — every call site treats a missing bean identically to a degraded call.
Client contract
com.attestpro.shared.llm.LlmClient:
| Method | Use |
|---|---|
chat(system, user, maxTokens) | Free-text completion |
chatStructured(system, user, maxTokens, jsonSchema) | JSON-Schema-constrained completion |
chatVisionStructured(system, prompt, images, mime, maxTokens, jsonSchema) | Image(s) + constrained JSON output |
embed(text) | Optional<double[]>, empty on failure |
Each feature owns a small PromptXxx class colocated with its service — system prompt,
user-prompt template, and JSON schema kept separate from orchestration logic.
Features
| Phase | Feature | Endpoint(s) |
|---|---|---|
| 1 | Verifier-facing insight — plain-language summary + cross-border equivalency note, included in the exported PDF report | POST /api/v1/credentials/insight |
| 2 | Admin dashboard anomaly narrative; plain-English → structured filter translation for audit-log search (model only ever produces a filter matching the page’s existing bounded param set, never executes a query itself) | GET /admin/network-stats/insight, POST /admin/audit-logs/nl-search |
| 3 | Bulk-issuance CSV column-mapping assistant (only when headers don’t already match); revocation-reason capture + classification into a fixed taxonomy; credential template field drafting | POST /portal/me/bulk-issuance/suggest-mapping, revoke flow, POST /portal/me/credential-templates/suggest-fields |
| 4 | AVETMISS field-mapping assistant — suggests scheme field-map values from a plain-language description, constrained to a fixed known list of expressions this platform can actually source | POST /admin/catalog/schemes/{schemeCode}/field-maps/suggest |
| 5 | Issuer KYC + holder ID document intelligence — vision extraction + cross-check, ephemeral (image discarded after one call) | POST /issuer/{did}/kyc-document/analyze, POST /holder/credentials/{id}/identity-check |
| 6 | Semantic catalog search (fallback behind exact/prefix search); docs/support assistant (retrieval-grounded, cites actual retrieved chunks) | GET /portal/catalog/entries/semantic-search, POST /support/ask |
Phase 3 detail
Every revocation classification attempt is persisted to core.llm_analysis_attempts
(feature=REVOCATION_CLASSIFY) — unlike the other Phase 3 features. Taxonomy:
MISCONDUCT / DATA_CORRECTION / EXPIRED_ACCREDITATION / ADMINISTRATIVE / OTHER.
Phase 5 detail
- A1 (
POST /api/v1/issuer/{did}/kyc-document/analyze, public, rate-limited 5/min — same trust level as registration since an issuer isn’t authenticated yet atPENDING_APPROVAL) +GET .../latest-analysis. Surfaces a loose name-match flag in the admin console — informational only, never gates approve/reject. - A2 (
POST /api/v1/holder/credentials/{id}/identity-check, holder-session-authenticated). Cross-checks an uploaded ID photo againstcredentialSubject.name(the only holder identity field the VC payload carries — DOB isn’t captured at issuance). Shown as “match confirmed” / “couldn’t confirm — contact your issuer”, never blocking any action. - Both endpoints enforce a 10MB upload cap.
Phase 6 detail
- Semantic catalog search: embeds the query, ranks a lazily-built, size-bounded, per-jurisdiction in-memory cache by cosine similarity. Known limitation: the cache only covers the first 1000 entries in default list order, not a query-relevant sample — acceptable because it’s a fallback behind primary text search, not the primary lookup.
- Docs/support assistant: on first request per process, chunks and embeds the documentation
tree (shipped into the jar at build time) — top-k cosine-similarity retrieval feeds a
structured-output call whose system prompt is constrained to answer only from retrieved
excerpts, instructed to say so rather than guess when they don’t cover the question. Returns
{answer, sources: [{doc, heading}]}— sources are the actually-retrieved chunks, so they can’t be hallucinated.