Files
Messe-Lotse/plan/ai-document-skills.md
T
2026-09-29 09:40:32 +00:00

1327 lines
64 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# AI Document Processing Skills — 24HRS Messe-Lotse
## Overview
Pluggable AI skills triggered on document upload to Pimcore DAM. Each skill extracts, validates, or enriches job data from uploaded files. All output is **non-destructive metadata** — the user explicitly applies or ignores each suggestion.
## Architecture
```
Asset uploaded to Pimcore DAM
│
▼ asset.postAdd event
DocumentClassifier
│
│ Determines type: briefing | brandbook | floorplan | power-confirmation |
│ wlan-confirmation | cleaning-confirmation | rigging-quote | rigging-drawing |
│ rigging-messe-confirmation | stand-approval | stand-photo | acceptance-photo |
│ weekly-protocol | other
│
▼ Routes to applicable extraction chain
Skill Pipeline (fan-out, parallel execution):
├── appropriate skills for this document type
│
▼
SuggestionService
│
│ Stores as JSON on job.aiSuggestions:
│ { id, skill, field, suggestedValue, confidence, sourceAssetId,
│ status: pending|applied|ignored, createdAt }
│
▼
Frontend SuggestionBadge
│
│ Rendered inline near affected form field
│ [Apply] → PATCH /api/messebau/jobs/{id}/suggestions/{suggestionId}/apply
│ [Ignore] → PATCH /api/messebau/jobs/{id}/suggestions/{suggestionId}/ignore
```
## Upload → Suggestion Workflow (Step by Step)
```
User: Drags file into "Briefing Kunde" upload area on Tab 1
│
▼
┌─────────────────────────────────────────────────────────────┐
│ STEP 1: Field Context (free) │
│ │
│ Upload field = briefingCustomerDoc │
│ → Context hint: "this is probably a customer briefing" │
│ → Context hint: belongs to Tab 1 (Briefing), not Tab 3 │
│ → Only brief/consultation skills will be triggered later │
└───────────────────────┬─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ STEP 2: Asset stored in Pimcore DAM │
│ │
│ - File saved, thumbnail generated │
│ - asset.postAdd event fires │
│ - AiSkillTriggerListener invoked (async, non-blocking) │
│ - Upload complete → UI shows success immediately │
└───────────────────────┬─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ STEP 3: DocumentClassifier + Mismatch Detection (always) │
│ │
│ DocumentClassifier analyzes: │
│ - File extension (.pdf, .dwg, .png) │
│ - MIME type │
│ - Filename ("Brevo-Zusammenfassung.pdf" → consultation) │
│ - First 4KB of extracted text ("Messeauftritte", │
│ "Briefing", "Halle 5.1" → confirm or refine type) │
│ │
│ Result: │
│ classifiedType = "consultation-summary" │
│ fieldContext = "briefingCustomerDoc" ✓ match │
│ confidence = 0.94 │
│ │
│ MismatchDetector checks: Does classifiedType fit this field? │
│ │
│ ✓ MATCH → proceed to field-specific skill chain (Step 5) │
│ ✗ MISMATCH (e.g. floor plan uploaded to briefing area) │
│ → STOP. Wait for user resolution before continuing: │
│ → [Move to Standplan] [Keep here anyway] [Cancel] │
└───────────────────────┬─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ STEP 4: Version Detection (if field has prior assets) │
│ │
│ DocumentVersionDetector compares new upload against │
│ existing assets in the same field: │
│ - Filename similarity (Levenshtein distance) │
│ - Content similarity (extracted dimensions, dates, text) │
│ │
│ NO prior version → skip, proceed to Step 5 │
│ SAME file detected → skip (duplicate), no skills triggered │
│ NEWER VERSION detected: │
│ → Diff extracted data (dimensions, budgets, dates) │
│ → If changes found: surface warning with severity │
│ → "Dimensions: 8.65×5.8m → 7.50×6.70m [Review] [Replace]"│
│ → Skills run on the NEW version only (old is superseded) │
│ DIFFERENT document (same name, different content): │
│ → "A different file with this name already exists" │
│ → [Keep both] [Replace] [Rename new] │
└───────────────────────┬─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ STEP 5: Field-Specific Skill Chain (parallel, async) │
│ │
│ Based on the upload FIELD (not just document type), │
│ only the relevant subset of skills is triggered. │
│ │
│ For upload field = briefingCustomerDoc: │
│ │
│ ┌─────────────────────┐ │
│ │ ContactExtractor │ → "Lisa Reinhardt" │
│ │ confidence: 0.92 │ → Tab 0 contact fields │
│ └─────────────────────┘ │
│ ┌─────────────────────┐ │
│ │ MultiFairDetector │ → 5 fairs detected │
│ │ confidence: 0.88 │ → modal: create fair records? │
│ └─────────────────────┘ │
│ ┌─────────────────────┐ │
│ │ BudgetExtractor │ → 30.000€, 5.000€ │
│ │ confidence: 0.95 │ → costEstimate field collection │
│ └─────────────────────┘ │
│ ┌─────────────────────┐ │
│ │ PainPointSummarizer │ → 5 pain points + 4 expectations │
│ │ confidence: 0.85 │ → job.notes suggestion │
│ └─────────────────────┘ │
│ ┌─────────────────────┐ │
│ │ ChecklistGenerator │ → 6 action items │
│ │ confidence: 0.82 │ → Tab 4 checklist │
│ └─────────────────────┘ │
│ ┌─────────────────────┐ │
│ │ DeadlineDetector │ → dates for 5 fairs │
│ │ confidence: 0.90 │ → Tab 0 setup/event/teardown │
│ └─────────────────────┘ │
│ ┌─────────────────────┐ │
│ │ MaterialListExtract │ → booth specs, furniture │
│ │ confidence: 0.78 │ → Tab 3 material sections │
│ └─────────────────────┘ │
│ │
│ Each runs independently. Fails gracefully (no cascade). │
│ Skills NOT in this field's matrix are never called. │
└───────────────────────┬─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ STEP 6: SuggestionService stores results │
│ │
│ job.aiSuggestions = │
│ [ │
│ { id: "uuid-1", skill: "ContactExtractor", │
│ field: "customer.contactName", │
│ suggestedValue: "Lisa Reinhardt", │
│ confidence: 0.92, sourceAssetId: 12345, │
│ status: "pending", createdAt: "2026-06-10T..." }, │
│ { id: "uuid-2", skill: "MultiFairDetector", │
│ field: "_fairs", │
│ suggestedValue: { fairs: [...] }, │
│ confidence: 0.88, status: "pending" }, │
│ { id: "uuid-3", skill: "BudgetExtractor", │
│ field: "costEstimate.standConstruction", │
│ suggestedValue: 30000, confidence: 0.95, │
│ status: "pending" }, │
│ ... │
│ ] │
└───────────────────────┬─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ STEP 7: UI Surfaces Suggestions │
│ │
│ Suggestions appear as badges across the app: │
│ │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ 📁 Brevo-Zusammenfassung.pdf ✅ │ │
│ │ │ │
│ │ 💡 4 AI suggestions ready (uploaded 2 min ago) │ │
│ │ │ │
│ │ Tab 0: Contact "Lisa Reinhardt" detected [Apply] │ │
│ │ Tab 0: Budget ~30.000€ detected [Apply] │ │
│ │ Fairs: 5 fairs detected → create records? [Review] │ │
│ │ Tab 4: 6 checklist items suggested [Review] │ │
│ │ Tab 3: 4 materials detected [Review] │ │
│ │ Notes: Pain point summary available [Preview] │ │
│ └─────────────────────────────────────────────────────────┘ │
│ │
│ Tab headers show pending counts: │
│ ┌──────────────────────────────────────────┐ │
│ │ 📋 Stammdaten (2) │ 📄 Briefing │ ... │ │
│ └──────────────────────────────────────────┘ │
│ │
│ User actions: │
│ - [Apply] → writes value to field, status = "applied" │
│ - [Apply All] (per tab) → bulk apply all in that tab │
│ - [Ignore] → status = "ignored", badge dismissed │
│ - [Preview] → expands to show full content before deciding │
│ - No action → stays as pending, visible on next page load │
│ - Suggestions auto-expire after 30 days (configurable) │
└─────────────────────────────────────────────────────────────┘
```
### When Does the User Need to Intervene?
| Scenario | What happens |
|---|---|
| Normal upload to correct field | Fully automatic. Suggestions surface passively. User reviews at their leisure. |
| Upload to wrong field (floor plan → briefing area) | MismatchDetector fires. **User must respond** before skills process (Move / Keep / Cancel). Skills are paused. |
| Duplicate file detected | VersionDetector finds identical file. No skills triggered. "This file was already uploaded on [date]" toast. |
| Newer version of existing doc | VersionDetector compares. User sees diff. User decides: Replace or Keep both. Skills run on the chosen version. |
| Low confidence suggestion (< AI_MIN_CONFIDENCE) | Suggestion stored but hidden by default. User can toggle "Show low confidence" to review. |
| LLM unavailable | Skills fall back to regex-only extraction. Confidence scores marked as `source: "regex"`. No user notification needed. |
| Skill throws exception | That skill fails silently. Other skills continue. Error logged. No impact on upload or other suggestions. |
---
## Field → Skill Routing Matrix
Each upload field in the UI maps to a specific subset of skills. Only these skills are triggered when a document is uploaded to that field. The common chain (Classifier, MismatchDetector, VersionDetector, **DocumentScopeDetector**) runs on every upload regardless of field.
Documents are stored at the entity level matching their natural scope:
```
customer.brandDoc ← company-wide brand guidelines
customer.consultationDocs ← multi-fair planning docs (e.g. Brevo consultation)
customer.contractDocs ← framework agreements
fair.constructionBriefing ← multi-booth briefings (e.g. CFC 2026, 112 booths)
fair.floorPlan ← hall-level plans, venue maps
fair.regulations ← venue regs, fire safety
fair.protocols ← fair-wide weekly protocols
job.briefingCustomerDoc ← per-job override (null → show customer.consultationDocs)
job.briefingPlanDoc ← per-job override (null → show fair.floorPlan)
job.briefingAdditionalDoc ← per-job additional docs
job.powerConfirmation etc. ← per-booth service confirmations
```
### Customer Detail — Dokumente Tab
| Upload field | Expected classifier type | Skills triggered | Notes |
|---|---|---|---|
| `customer.brandDoc` | `brandbook` | `BrandConsistencyChecker` (preloads palette) | No immediate suggestion. Palettes used for later photo checks |
| `customer.consultationDocs` | `consultation-summary`, `briefing` | `ContactExtractor`, `MultiFairDetector`, `BudgetExtractor`, `PainPointSummarizer`, `ChecklistGenerator`, `DeadlineDetector` | Heavy chain — same as briefingCustomerDoc was |
| `customer.contractDocs` | `other` (generic) | `DeadlineDetector`, `BudgetExtractor` | Contract terms, payment deadlines |
### Fair Detail — Dokumente Tab
| Upload field | Expected classifier type | Skills triggered | Notes |
|---|---|---|---|
| `fair.constructionBriefing` | `briefing`, `floorplan` | `StandDataExtractor`, `MaterialListExtractor`, `DeadlineDetector`, `ChecklistGenerator` | Multi-booth briefings → suggestions scoped to fair, propagated to all jobs |
| `fair.floorPlan` | `floorplan` | `StandDataExtractor`, `DeadlineDetector` | Hall plans, venue maps |
| `fair.regulations` | `other` | `DeadlineDetector` | Fire safety, venue rules → deadline extraction |
| `fair.protocols` | `weekly-protocol` | `WeeklyProtocolSummarizer` | Fair-wide protocols |
### Tab 0 — Stammdaten & Termine
| Upload field | Expected classifier type | Skills triggered | Notes |
|---|---|---|---|
| `standImage` | `stand-photo` | None (first upload). On 2nd+ upload: `SetupProgressAnalyzer`, `BrandConsistencyChecker` | BrandConsistencyChecker only runs if brandbook was previously uploaded to `briefingBrandDoc` |
### Tab 1 — Briefing (Job Merged View)
The job's Tab 1 shows documents from all three scope levels merged. Upload fields are per-level:
| Upload field | Entity | Expected classifier type | Skills triggered | Notes |
|---|---|---|---|---|
| `customer.brandDoc` (via Briefing tab) | customer | `brandbook` | `BrandConsistencyChecker` | Uploaded through the "☰ Kunde" section |
| `customer.consultationDocs` (via Briefing tab) | customer | `consultation-summary`, `briefing` | `ContactExtractor`, `MultiFairDetector`, `BudgetExtractor`, `PainPointSummarizer`, `ChecklistGenerator`, `DeadlineDetector` | Uploaded through the "☰ Kunde" section |
| `fair.constructionBriefing` (via Briefing tab) | fair | `briefing`, `floorplan` | `StandDataExtractor`, `MaterialListExtractor`, `DeadlineDetector`, `ChecklistGenerator` | Uploaded through the "🌐 Messe" section; multi-booth scope |
| `fair.floorPlan` (via Briefing tab) | fair | `floorplan` | `StandDataExtractor`, `DeadlineDetector` | Uploaded through the "🌐 Messe" section |
| `fair.regulations` (via Briefing tab) | fair | `other` | `DeadlineDetector` | Uploaded through the "🌐 Messe" section |
| `job.briefingCustomerDoc` (override) | job | `briefing` | `DeadlineDetector` | Per-job override; if null, inherits from customer |
| `job.briefingPlanDoc` (override) | job | `floorplan` | `StandDataExtractor`, `DeadlineDetector` | Per-job override; if null, inherits from fair |
| `job.briefingAdditionalDoc` | job | `other` (generic) | `DeadlineDetector` | Minimal chain — scan for dates only |
### Tab 2 — Organisation
| Upload field | Expected classifier type | Skills triggered | Notes |
|---|---|---|---|
| `powerConfirmation` | `confirmation` | `ConfirmationParser`, `DeadlineDetector`, `BudgetExtractor` | ConfirmationParser compares booked vs actual service details |
| `wlanConfirmation` | `confirmation` | `ConfirmationParser`, `DeadlineDetector`, `BudgetExtractor` | — |
| `cleaningConfirmation` | `confirmation` | `ConfirmationParser`, `DeadlineDetector` | — |
| `riggingQuote` | `confirmation` (quote subtype) | `BudgetExtractor`, `DeadlineDetector` | Quote subtype triggers budget comparison |
| `riggingDrawing` | `floorplan` | `StandDataExtractor` | Technical drawing — extract rigging points, dimensions |
| `riggingMesseConfirmation` | `confirmation` | `ConfirmationParser`, `DeadlineDetector` | — |
| `standApprovalDoc` | `confirmation` | `ConfirmationParser`, `DeadlineDetector` | — |
### Tab 5 — Standfotos
| Upload field | Expected classifier type | Skills triggered | Notes |
|---|---|---|---|
| `standGallery` (1st photo) | `stand-photo` | None | Photos stored for gallery and later analysis |
| `standGallery` (2nd+ photo) | `stand-photo` | `SetupProgressAnalyzer`, `BrandConsistencyChecker` | Compares latest two photos for structural changes; checks brand consistency if brandbook loaded |
### Tab 6 — Standabnahme
| Upload field | Expected classifier type | Skills triggered | Notes |
|---|---|---|---|
| `acceptancePhotos` | `acceptance-photo` | None (stored for later) | Photos are retrieved when user clicks "Generate Acceptance Report" (manual trigger for `AcceptanceReportGenerator`) |
| `acceptanceSignature` | `signature-image` | None | Signature is canvas-generated, not document-uploaded |
### Tab 7 — Weekly
| Upload field | Expected classifier type | Skills triggered | Notes |
|---|---|---|---|
| `weeklyProtocols` | `weekly-protocol` | `WeeklyProtocolSummarizer` | Generates structured summary card from each protocol PDF |
---
## Common Chain (Runs on Every Upload, Before Skill Chain)
| Step | Service | Always runs? | Blocking? |
|---|---|---|---|
| 1. Classify | `DocumentClassifier` | Yes, every upload | No — async, result enriches subsequent steps |
| 2. Mismatch check | `DocumentMismatchDetector` | Yes, every upload | **Yes** — if mismatch detected, skills are paused until user resolves |
| 3. Version check | `DocumentVersionDetector` | Yes, if field has ≥1 prior asset | No — runs in parallel with skills; surfaces warning if conflict found |
| 4. Skill chain | Field-specific (see matrix above) | Only if mismatch is resolved AND skills are defined for this field | No — runs async, suggestions surface progressively |
---
## Skill Catalog
### Tier 1 — High Impact (saves manual data entry)
---
#### SKILL-001: DocumentClassifier
| Property | Value |
|---|---|
| **Trigger** | Every `asset.postAdd` |
| **Input** | File metadata (extension, MIME type, filename), first 4KB of content |
| **Output** | Document type enum + confidence |
| **Classification rules** | |
| Pattern | Type |
|---|---|
| `*.pdf` in `briefing-*` fields | `briefing` |
| `*.ai`, `*.eps`, `*.svg` in `briefing-brand-*` | `brandbook` |
| `*.dwg`, `*.dxf`, `*grundriss*`, `*plan*`, `*zeichnung*` | `floorplan` |
| `*bestätigung*`, `*confirmation*`, `*auftrag*` | `confirmation` |
| `*.jpg`, `*.jpeg`, `*.png` in `standImage` or `standGallery` | `stand-photo` |
| `*.jpg`, `*.jpeg`, `*.png` in `acceptancePhotos` | `acceptance-photo` |
| `*.pdf` in `weeklyProtocols` | `weekly-protocol` |
| Content contains "Angebot", "Kostenvoranschlag" | `confirmation` (quote subtype) |
| Unknown | `other` |
**No LLM needed.** Pure regex + metadata matching.
---
#### SKILL-001a: DocumentScopeDetector
| Property | Value |
|---|---|
| **Trigger** | Every `asset.postAdd` on a briefing-type field (runs during classification Step 3) |
| **Input** | Extracted text + upload field context + current entity's parent (job's customer + fair) |
| **Output** | `{ suggestedScope: "customer"\|"fair"\|"job", suggestedEntityId, reason, mismatch: boolean }` |
**Detection heuristics:**
| Document content pattern | Suggested scope | Example |
|---|---|---|
| Mentions multiple fairs by name (≥2) | `customer` | "OMR, Dmexco, K5, E-Commerce Berlin, Retouren-Messe" |
| Contains "Brandbook", "Corporate Identity", "Logo", "Style Guide" | `customer` | Brandbook applies company-wide |
| Contains "Rahmenvertrag", "AGB", "Framework Agreement" | `customer` | Contract terms |
| Mentions "alle Stände", "X Stände", "insgesamt X", "Gesamt", booth count > 5 | `fair` | "112 Messestände in Halle 5.1" |
| Contains "Halle X" with booth type breakdown (multiple sizes) | `fair` | CFC doc: 88×4×2m, 20×5×5m, 4×10×5m |
| References "Hallenplan", "Venue Map", "Messe [City] Vorschriften" | `fair` | Venue regulations |
| References one specific booth + one company | `job` | "Brevo, 50m², Hall A4" |
| Contains "Standfoto", "Abnahme", booth photo metadata | `job` | Photos are always job-scoped |
**Uploaded to wrong level:**
```
User uploads CFC briefing to job.briefingCustomerDoc
→ Classifier: "briefing", DocumentScopeDetector: scope=fair, 112 booths
→ MISMATCH: field expects job-scoped, document is fair-scoped
UI shows:
┌──────────────────────────────────────────────────────────────┐
│ ⚠️ This document appears to be fair-scoped, not job-scoped. │
│ (references 112 booths in Halle 5.1). │
│ │
│ [Move to fair: "Cashflow Conference 2026"] │
│ [Keep here (this is a job-specific excerpt)] │
│ [Cancel] │
└──────────────────────────────────────────────────────────────┘
```
**If user moves:** Asset relocated to `fair.constructionBriefing`. Skills re-run on the FAIR entity. Results are fair-scoped with "Apply to all 112 jobs?" propagation.
**If user keeps:** Skills run on the job entity. Warning badge remains. Skills may produce false positives (e.g. MultiFairDetector finding 5 fairs in a doc kept on one job).
**Confidence:** 0.90+ for clear patterns (multi-fair, booth counts, "alle Stände"). 0.70 for ambiguous cases (single booth but with venue context words).
---
#### SKILL-002: ContactExtractor
| Property | Value |
|---|---|
| **Trigger** | Document classified as `briefing`, `confirmation` |
| **Input** | Extracted text from PDF/DOCX |
| **Output** | `{ name, email, phone, role }[]` with confidence per field |
| **Target field** | `customer.contactName`, `customer.contactPhone`, `customer.contactEmail` |
| **Suggestion UI** | Tab 0, Ansprechpartner section |
**Extraction strategy** (ordered by preference):
1. **Regex** — Email: `[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}`, Phone: `(?:\+49|0)[\s.-]?[1-9]\d{1,4}[\s.-]?\d{3,}[\s.-]?\d{3,}`
2. **LLM prompt** (if regex misses): "Extract all contacts from this document. Return JSON array of {name, email, phone, role}."
3. **Comparison**: If extracted values differ from stored `customer` contact fields, suggestion shows both values with yellow "Override?" flag.
**Confidence scoring:**
- Email regex match: 0.95
- Phone regex match: 0.85
- Name from LLM: 0.70
- Role from LLM: 0.50
---
#### SKILL-003: DeadlineDetector
| Property | Value |
|---|---|
| **Trigger** | Any document (all types) |
| **Input** | Extracted text |
| **Output** | `{ date, recognisedHoliday, suggestedField, urgency }[]` |
| **Target fields** | All deadline fields across Tab 0/2/3 |
**Date patterns recognized:**
- `DD.MM.YYYY`, `DD.MM.YY`
- `YYYY-MM-DD`
- `DD. Month YYYY`
- German: `bis zum 15. Mai 2026`, `spätestens 30.06.2026`
- Relative: `in 2 Wochen`, `nächsten Montag` (resolved against document metadata date)
**Field mapping heuristics:**
| Text contains | Maps to field |
|---|---|
| `strom`, `power`, `kW`, `elektro` | `powerBooking.deadline` |
| `wlan`, `wifi`, `internet` | `wlan.deadline` |
| `reinigung`, `cleaning` | `cleaning.deadline` |
| `rigging`, `truss`, `traverse` | `rigging.deadline` |
| `standfreigabe`, `genehmigung` | `standApproval.deadline` |
| `druck`, `print`, `produktion` | `druckproduktion.deadline` |
| `bestellung`, `booking`, `buchung` | `bookingPlatform.deadline` |
**Conflict detection**: If document date differs from stored deadline, flag as "mismatch" with amber warning badge.
---
#### SKILL-004: ChecklistGenerator
| Property | Value |
|---|---|
| **Trigger** | Document classified as `briefing` |
| **Input** | Extracted text |
| **Output** | `{ text, confidence, sourceQuote }[]` |
| **Target field** | `checklist[]` on job |
| **Suggestion UI** | Tab 4, "Add suggested items" button above checklist |
**Extraction strategy:**
1. **LLM prompt**: "You are analyzing a trade fair booth briefing document. Extract all actionable requirements, specifications, and tasks. For each, provide: a short task description in German, a confidence score (0-1), and the source sentence from the document. Return as JSON array."
2. **Deduplication** against existing checklist items by cosine similarity of text embeddings
**Example output:**
```json
[
{"text": "2 Banner 3×2m mit Ösen produzieren", "confidence": 0.92, "source": "Wir benötigen zwei Banner im Format 3×2m mit Ösen"},
{"text": "Teppichboden in Grau verlegen", "confidence": 0.88, "source": "Der Boden soll mit grauem Teppich ausgelegt werden"}
]
```
---
#### SKILL-004a: DeadlineDetector — German Date Format Enhancements
| Property | Value |
|---|---|
| **Trigger** | Any document (all types) — extends SKILL-003 |
| **Enhancement** | German date range patterns found in real trade fair documents |
**Additional patterns recognized:**
- `Aufbau: 01.–02.07.2026` → two dates: setupStart=2026-07-01, setupEnd=2026-07-02
- `Event: 03.–04.07.2026` → two dates: eventStart=2026-07-03, eventEnd=2026-07-04
- `Abbau ab 04.07. 16:30 Uhr` → teardownStart=2026-07-04 16:30:00, teardownEnd=2026-07-04T23:59:59
- `DD.–DD. Month YYYY` → German date range spanning two days
- `Halle` + number → maps to `hall` field as "Halle X"
- `Messe [City]` → maps to fairLocation name inference
- Event name detection: "Cashflow Conference 2026", "OMR Festival 2026" → maps to `fair.name`
- `Standhöhe X m` → maps to `standHeight`
- `Traversenhöhe X m` → stored as notes
**Field mapping for German context:**
| Text contains | Maps to field |
|---|---|
| `Aufbau:` | `setupStart`, `setupEnd` |
| `Event:`, `Messe:` | `eventStart`, `eventEnd` |
| `Abbau` | `teardownStart`, `teardownEnd` |
---
#### SKILL-004b: MultiFairDetector
| Property | Value |
|---|---|
| **Trigger** | Document classified as `briefing` or `consultation-summary` |
| **Input** | Extracted text |
| **Output** | `{ fairs: [{ name, size, standType, locationHint, estimatedBudget }], action: "createFairs" }` |
| **Target UI** | Modal: "This document references 5 fairs. Create these as fair records for customer {customer.company}?" |
**Extraction strategy:**
1. **LLM prompt**: "Extract all trade fair / event references from this document. For each: fair name, booth size in m², stand type if mentioned, location/hall, budget if mentioned. Return JSON array."
2. **Deduplication**: Compare against existing fairs in the system (by name + year) — skip already existing
3. **Auto-link**: If document references a `customer`, pre-select that customer for all created fairs
**Example from Brevo consultation doc:**
```json
{
"fairs": [
{"name": "OMR Festival 2026", "size": 50, "standType": "Eckstand", "estimatedBudget": 40000},
{"name": "Dmexco 2026", "size": 40, "standType": "Eckstand", "estimatedBudget": 30000},
{"name": "K5 2026", "size": null, "estimatedBudget": null},
{"name": "E-Commerce Berlin Expo 2026", "size": 18, "estimatedBudget": null},
{"name": "Retouren-Messe 2026", "size": 15, "estimatedBudget": null}
],
"action": "createFairsWithJobs",
"customerHint": "Brevo"
}
```
---
#### SKILL-004c: BudgetExtractor
| Property | Value |
|---|---|
| **Trigger** | Any document containing currency amounts |
| **Input** | Extracted text |
| **Output** | `{ amounts: [{ value, currency, context, suggestedField }], totalDetected }` |
| **New field** | `job.costEstimate` field collection with fields: `standConstruction`, `activation`, `furniture`, `logistics` (all integer, currency EUR) |
**Extraction strategy:**
1. **Regex**: `(?:Budget|Kosten|ca\.?|approx\.?)?\s*(\d{1,3}(?:[.,]\d{3})*(?:[.,]\d+)?)\s*(?:€|EUR|Euro|k€)` with lookbehind context
2. **Context mapping**:
| Surrounding text | Maps to |
|---|---|
| `Standbau`, `Booth construction`, `Bau` | `costEstimate.standConstruction` |
| `Aktivierung`, `Activation`, `Lead Magnet` | `costEstimate.activation` |
| `Möbel`, `Mobiliar`, `Furniture` | `costEstimate.furniture` |
| `Transport`, `Logistik` | `costEstimate.logistics` |
| `k€` | multiply by 1000 |
**Example from Brevo OMR doc:**
```
standConstruction: 40.000€ (source: "Budget Approx. 40k€")
activation: 5.000€ (source: "Budget approx. 5k€")
→ Suggestion: "Apply these budget estimates to the job?"
```
**Example from CFC 2026 doc:**
```
→ "Multi-booth briefing (112 booths). Create costEstimate per booth type?"
```
---
#### SKILL-004d: PainPointSummarizer
| Property | Value |
|---|---|
| **Trigger** | Document classified as `consultation-summary` or `briefing` |
| **Input** | Extracted text |
| **Output** | `{ painPoints: string[], expectations: string[], structuredNotes: string }` |
| **Target field** | `job.notes` (prepopulated or appended) |
**Extraction strategy:**
1. **LLM prompt**: "This is a consultation summary for a trade fair booth project. Extract: 1) Pain points / challenges as short bullet points in German, 2) Customer expectations as bullet points in German, 3) A concise structured notes summary (3-5 sentences) in German. Return JSON."
2. **Suggestion UI**: "AI extracted 5 pain points, 4 expectations, and a structured summary. Append to job notes? [Preview] [Append] [Replace]"
**Example from Brevo consultation:**
```json
{
"painPoints": [
"Wiederholte Turnkey-Miete teurer als Kauf",
"Unterschiedliche Standdesigns verhindern einheitlichen Markenauftritt",
"Standard-Messeauftritt reicht nicht zur Wettbewerbsdifferenzierung",
"Ineffektive Lead-Generierung, Partner soll proaktiv Ideen einbringen",
"Mietstrategie nicht flexibel für variierende Standgrößen (10-50m²)"
],
"expectations": [
"Modulares Standdesign, skalierbar 10-50m²",
"Strategische Unterscheidung OMR (Prestige) vs Dmexco (Sales)",
"Kreative Aktivierungsideen passend zum 'All-in-one'-Gedanken",
"Proaktiver Partner für Formalitäten und transparente Kommunikation"
],
"structuredNotes": "Kunde Brevo plant 5 Messeauftritte 2026 (OMR 50m², Dmexco 32-50m², K5, E-Commerce Berlin Expo, Retouren-Messe). Budget ca. 30.000€ pro Großmesse. Wünscht modulares, skalierbares Standdesign und kreative Lead-Generierung. Kontakt: Lisa Reinhardt."
}
```
---
### Tier 2 — Validation (compares documents against stored data)
---
#### SKILL-005: StandDataExtractor
| Property | Value |
|---|---|
| **Trigger** | Document classified as `floorplan` |
| **Input** | DWG/DXF/PDF rendering or text layer |
| **Output** | `{ width, depth, standType, hall, standNumber, area }` |
| **Target fields** | `standSize`, `standType`, `hall`, `area` on Tab 0 |
**Extraction approach:**
- **PDF/DWG with text layer**: Parse legend/description block for dimensions
- **Image-based PDF**: Use LLM vision model: "Extract the booth dimensions, stand type (from legend), hall number, and stand number from this floor plan."
- **Comparison**: If extracted dimensions differ >10% from stored `standSize`, flag as discrepancy
---
#### SKILL-006: ConfirmationParser
| Property | Value |
|---|---|
| **Trigger** | Document classified as `confirmation` |
| **Input** | Extracted text from booking confirmation PDF |
| **Output** | `{ serviceType, details, confirmationNumber, bookingDate, status: confirmed|pending|rejected }` |
| **Target fields** | `powerBooking`, `wlan`, `cleaning`, `rigging`, `standApproval` |
**Extraction:**
- **LLM prompt**: "Extract from this booking confirmation: service type (power/WLAN/cleaning/rigging/stand approval), service details (kW, type, etc.), confirmation number, booking date, status (confirmed/pending/rejected). Return JSON."
- **Auto-confirm**: If `powerBooking.confirmationPdf` exists AND confirmation number is found in text → set `statusConfirmed` flag on the booking
- **Discrepancy flag**: If booked 6kW but `job.powerBooking.kw = 3.5` → warn
---
#### SKILL-007: BrandConsistencyChecker
| Property | Value |
|---|---|
| **Trigger** | Stand photo uploaded AND brandbook exists for this job |
| **Input** | Brandbook (AI/EPS/PDF) + stand photo (JPG/PNG) |
| **Output** | `{ colorMatches: bool, primaryColorDelta, logoVisible: bool, logoPosition: string, issues[] }` |
| **Target UI** | Notification on Tab 5 Standfotos |
**Extraction:**
1. Extract brand colors from brandbook (dominant palette)
2. Analyze stand photo pixels for color histogram
3. Delta-E comparison between brand primary and photo dominant colors
4. **LLM vision**: "Is the company logo visible in this trade fair booth photo? If yes, where is it positioned? Are there any deviations from the brand guidelines visible?"
5. Result: green checkmark or amber warning in photo gallery overlay
---
#### SKILL-008: DocumentMismatchDetector
| Property | Value |
|---|---|
| **Trigger** | Any asset upload |
| **Input** | File metadata + classified type vs upload target field |
| **Output** | `{ suggestedField, reason }` |
| **Suggestion UI** | Banner: "This looks like a floor plan, but was uploaded to Briefing Kunde. Move to Standplan?" |
**Rules:**
| Uploaded to | Classified as | Suggestion |
|---|---|---|
| `briefingCustomerDoc` | `floorplan` | Move to `briefingPlanDoc` |
| `briefingCustomerDoc` | `brandbook` | Move to `briefingBrandDoc` |
| `powerConfirmation` | `wlan-confirmation` | Move to `wlanConfirmation` |
| `standGallery` | `acceptance-photo` | Move to `acceptancePhotos` |
---
#### SKILL-008a: DocumentVersionDetector
| Property | Value |
|---|---|
| **Trigger** | New asset uploaded to a field that already has an existing asset |
| **Input** | New upload + existing assets in same field |
| **Output** | `{ isNewerVersion: boolean, previousAssetId, changes: string[], severity: info|warning|critical }` |
| **Target UI** | Warning banner above the asset field: "A previous version exists. Changed: dimensions 8.65×5.8m → 7.50×6.70m. [Compare] [Keep both] [Replace]" |
**Detection strategy:**
1. **Filename similarity**: Levenshtein distance < 5 and shared substrings (>60% match) → likely same document, different version
2. **Content comparison**: Extract structured data (dimensions, budgets, dates) from both versions via same extraction pipeline → diff the outputs
3. **Timestamp**: Newer metadata date → newer version
**Real example from sample docs (OMR 2026):**
```
Previous version: OMR 2026_Booth Details & Ideation.pdf
→ dimensions: 8.65 × 5.80m, area: 50.25m²
New version: OMR 2026_Booth Details.pdf
→ dimensions: 7.50 × 6.70m, area: 50.25m²
Detected changes (severity: warning):
- Width: 8.65m → 7.50m (-13%)
- Depth: 5.80m → 6.70m (+16%)
- Area unchanged: 50.25m²
Suggestion: "Booth dimensions changed significantly but area is the same. Review booth layout? Update job?"
```
**Severity thresholds:**
- `info`: File replaced with same dimensions
- `warning`: Dimensions changed but area within 10%
- `critical`: Area changed >10%, or key requirements changed
---
### Tier 3 — Enhancement (adds convenience)
---
#### SKILL-009: MaterialListExtractor
| Property | Value |
|---|---|
| **Trigger** | Document classified as `briefing` or `floorplan` |
| **Input** | Extracted text |
| **Output** | `{ materialType: flooring|standsystem|hardware|furniture, name, suggestedQuantity, confidence }[]` |
| **Target UI** | Tab 3, "Add suggested materials" button per material section |
**LLM prompt**: "Extract all materials, furniture, hardware, and equipment mentioned in this document. Categorize each as: flooring (Bodenbeläge), standsystem (stand elements), hardware (technical equipment), or furniture (Möbel). Return JSON with name, category, and suggested quantity."
---
#### SKILL-009a: LeadMagnetIdeator
| Property | Value |
|---|---|
| **Trigger** | Document classified as `briefing` or `consultation-summary` with activation/lead magnet requirements |
| **Input** | Full extracted text including company description, product info, budget, previous activations |
| **Output** | `{ ideas: [{ title, description, relevance, estimatedBudget, difficulty }] }` |
| **Target UI** | Side panel on Tab 3 Leadmagneten section: "AI generated 3 activation ideas based on this briefing →" |
**LLM prompt**: "You are a creative trade fair booth designer. The following document describes a customer, their product, booth requirements, and budget for activation/lead magnets. Generate 3 creative lead magnet / booth activation ideas that: 1) tie directly to the company's product or service, 2) fit the specified budget, 3) are appropriate for the event type and audience, 4) drive measurable lead generation. Return JSON array with title, description (2-3 sentences), relevance score (0-1), estimated cost range, and implementation difficulty (easy/medium/hard)."
**Example from Brevo OMR document:**
```json
{
"ideas": [
{
"title": "AI Marketing Campaign Builder Challenge",
"description": "Visitors build a real marketing campaign using Brevo's platform in 60 seconds. Best campaign wins a premium Brevo subscription. Showcases the 'all-in-one' platform capability live. Data capture via campaign creation form doubles as lead gen.",
"relevance": 0.92,
"estimatedBudget": "3.000-5.000€",
"difficulty": "medium"
},
{
"title": "Omnichannel Journey Maze",
"description": "Physical maze where visitors must choose the right channel (email/SMS/WhatsApp/push) at each decision point to reach the customer. Screens show Brevo automation workflows. Gamified lead capture with scoreboard and daily prizes.",
"relevance": 0.85,
"estimatedBudget": "5.000-8.000€",
"difficulty": "hard"
},
{
"title": "Brevo Customer Data Wall",
"description": "Interactive touch wall where visitors explore anonymized customer journey data. 'Find the pattern' challenge with instant Brevo AI analysis. Captures email for results delivery. Instagram-worthy data visualization backdrop.",
"relevance": 0.78,
"estimatedBudget": "3.000-6.000€",
"difficulty": "medium"
}
]
}
```
**Suggestion UI**: Ideas shown as cards with title, description, difficulty badge. User can select one and "Add to lead magnets" which populates the notes section of the lead magnets block on Tab 3.
**Context for better ideas**: The skill consumes:
- Company description (from the document)
- Product information (from the document)
- Previous successful activations (from document or previous jobs for this customer)
- Budget constraints (from BudgetExtractor if run first)
- Event type and atmosphere (from document)
---
#### SKILL-010: AcceptanceReportGenerator
| Property | Value |
|---|---|
| **Trigger** | Manual: "Generate Acceptance Report" button on Tab 6 |
| **Input** | `acceptancePhotos[]`, `acceptanceSignature`, `acceptanceRemarks`, `acceptanceDate`, job metadata |
| **Output** | PDF via Gotenberg |
| **Target UI** | Download link on Tab 6 |
**LLM prompt**: "You are generating a trade fair booth acceptance report. Analyze the acceptance photos and remarks. Create a structured report in German with: 1) Header with job info, 2) Photo gallery with annotations of any visible issues, 3) Remarks section, 4) Issue summary, 5) Sign-off block with signature and date."
The Twig template renders the LLM output as a formatted PDF with photos embedded inline.
---
#### SKILL-011: SetupProgressAnalyzer
| Property | Value |
|---|---|
| **Trigger** | New stand photo uploaded, AND at least one prior photo exists in `standGallery` |
| **Input** | Two most recent stand photos (chronological) |
| **Output** | `{ detectedChanges[], autoCheckedChecklistItems[] }` |
| **Target UI** | Toast: "Build progress detected: walls erected. Checklist item 'Standaufbau-Genehmigung einholen' was checked." |
**LLM vision prompt**: "Compare these two photos of a trade fair booth under construction. Photo A is older, Photo B is newer. What structural changes are visible? Options: walls erected, flooring laid, graphics/panels mounted, furniture placed, lighting installed, equipment installed, cleaning completed."
**Auto-check mapping:**
| Detected change | Auto-checked checklist item |
|---|---|
| `walls erected` | — (manual review) |
| `graphics mounted` | "Grafiken zur Produktion freigeben" |
| `furniture placed` | "Material & Werkzeug einpacken" |
| `equipment installed` | "Technische Anschlüsse bestellen" |
| `flooring laid` | — (manual review) |
All auto-checks have low confidence (0.75) and can be manually toggled back.
---
#### SKILL-012: WeeklyProtocolSummarizer
| Property | Value |
|---|---|
| **Trigger** | New PDF uploaded to `weeklyProtocols` |
| **Input** | Extracted text from protocol PDF |
| **Output** | `{ date, statusUpdate, openItems[], decisions[], nextSteps }` |
| **Target UI** | Structured summary card on Tab 7 Weekly, above the file list |
**LLM prompt**: "Summarize this weekly meeting protocol in German. Extract: 1) Date of meeting, 2) Status update (1-2 sentences), 3) Open items (array of strings), 4) Decisions made (array of strings), 5) Next steps (array of strings). Return JSON."
**Rendering**: Summary cards accumulate chronologically, building a project timeline visible on the Weekly tab.
---
## Skill Execution Architecture
### Overview
Skills are invoked via Symfony Messenger — already bundled with Pimcore. The listener dispatches one async message per applicable skill. Multiple worker processes consume the queue independently, achieving parallel execution without threads. No new infrastructure required.
### Why Not n8n
n8n adds a separate execution runtime, network latency (3 HTTP hops per skill: Pimcore → n8n → LLM → n8n → Pimcore), and an operational dependency. For this use case — "receive event → extract text → call LLM → store result" — Symfony Messenger is simpler, faster, and already in the stack (the same transport used for email notifications in TASK-020).
### Execution Flow
```
Pimcore asset.postAdd
│
▼
AiSkillTriggerListener (sync, < 5ms)
│
│ Classifier runs inline (regex-only, fast)
│ MismatchDetector runs inline (must resolve before skills proceed)
│ If mismatch → return immediately, skill chain paused
│ If match → dispatch one AnalyzeDocumentMessage per applicable skill
│
│ MessageBus::dispatch(new AnalyzeDocumentMessage(
│ assetId: 12345,
│ jobId: 42,
│ skillClass: ContactExtractor::class,
│ documentType: "consultation-summary"
│ ));
│ // ...dispatched for each skill in the field's routing matrix
│
▼
┌──────────────────────────────────────────────────────────┐
│ Symfony Messenger Transport │
│ (Redis or Doctrine) │
│ │
│ Queue: ai_skills │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ AnalyzeDoc │ │ AnalyzeDoc │ │ AnalyzeDoc │ │
│ │ ContactEx │ │ MultiFairDet │ │ BudgetEx │ ... │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
└─────────┼────────────────┼────────────────┼──────────────┘
│ │ │
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│ Worker 1 │ │ Worker 2 │ │ Worker 3 │ ← separate PHP processes
│ (proc) │ │ (proc) │ │ (proc) │
│ │ │ │ │ │
│ Extract │ │ Extract │ │ Extract │
│ text → │ │ text → │ │ text → │
│ regex → │ │ regex → │ │ regex → │
│ LLM fall │ │ LLM fall │ │ LLM fall │
│ → Svc │ │ → Svc │ │ → Svc │
└──────────┘ └──────────┘ └──────────┘
```
**How "parallel" works in PHP:** Not threads. Multiple `bin/console messenger:consume` processes. Each is an independent PHP process pulling from the same queue. Run 4 workers → 7 skills complete in roughly the time of the slowest skill, not the sum.
### Code: Listener
```php
// bundles/MessebauBundle/EventListener/AiSkillTriggerListener.php
class AiSkillTriggerListener
{
public function __construct(
private DocumentClassifier $classifier,
private DocumentMismatchDetector $mismatchDetector,
private FieldSkillRoutingMatrix $routingMatrix,
private MessageBusInterface $messageBus,
) {}
public function onAssetAdd(AssetEvent $event): void
{
$asset = $event->getAsset();
$job = $this->resolveJobFromAsset($asset);
if (!$job) return; // Asset not linked to a job
// Step 1: Classify inline (fast, regex-only, no LLM)
$classifiedType = $this->classifier->classify($asset);
// Step 2: Mismatch check (sync — must resolve before skills)
if ($mismatch = $this->mismatchDetector->detect($asset, $classifiedType)) {
$this->suggestionService->storeMismatch($job->getId(), $asset, $mismatch);
return; // Skills paused until user resolves
}
// Step 3: Dispatch one message per applicable skill
$fieldName = $asset->getField(); // e.g. "briefingCustomerDoc"
$skillClasses = $this->routingMatrix->getSkillsForField($fieldName);
foreach ($skillClasses as $skillClass) {
$this->messageBus->dispatch(new AnalyzeDocumentMessage(
assetId: $asset->getId(),
jobId: $job->getId(),
skillClass: $skillClass,
documentType: $classifiedType,
));
}
}
private function resolveJobFromAsset(Asset $asset): ?Job
{
// Walk Pimcore relations: find the Job that references this asset
// Pimcore's DependencyService provides a reverse lookup
$dependencies = \Pimcore\Model\Dependency::getBySourceId(
$asset->getId(), 'asset'
);
foreach ($dependencies->getRequiredBy() as $dep) {
if ($dep['type'] === 'object' && str_starts_with($dep['subtype'], 'Job')) {
return Job::getById($dep['id']);
}
}
return null;
}
}
```
### Code: Message + Handler
```php
// bundles/MessebauBundle/Message/AnalyzeDocumentMessage.php
class AnalyzeDocumentMessage
{
public function __construct(
public readonly int $assetId,
public readonly int $jobId,
public readonly string $skillClass,
public readonly string $documentType,
) {}
}
// bundles/MessebauBundle/MessageHandler/AnalyzeDocumentHandler.php
class AnalyzeDocumentHandler implements MessageHandlerInterface
{
public function __construct(
private SkillRegistry $skillRegistry,
private SuggestionService $suggestionService,
private array $config,
) {}
#[AsMessageHandler]
public function __invoke(AnalyzeDocumentMessage $message): void
{
$asset = Asset::getById($message->assetId);
if (!$asset) return; // Asset deleted before message processed
// Instantiate the skill via registry
$skill = $this->skillRegistry->get($message->skillClass);
// Run extraction (regex → LLM fallback if needed)
$result = $skill->analyze($asset, $message->documentType);
// Store if confidence meets threshold
if ($result->confidence >= ($this->config['AI_MIN_CONFIDENCE'] ?? 0.70)) {
$this->suggestionService->store(
jobId: $message->jobId,
skill: $message->skillClass,
field: $result->targetField,
suggestedValue: $result->value,
confidence: $result->confidence,
sourceAssetId: $message->assetId,
sourceAssetName: $asset->getFilename(),
);
}
}
}
```
### Code: Skill Interface + Registry
```php
// bundles/MessebauBundle/Service/AiSkill/AiSkillInterface.php
interface AiSkillInterface
{
public function analyze(Asset $asset, string $documentType): SkillResult;
}
// SkillResult is a value object:
// { targetField: string, value: mixed, confidence: float, source: "regex"|"llm", metadata: array }
// bundles/MessebauBundle/Service/AiSkill/SkillRegistry.php
class SkillRegistry
{
/** @var array<string, AiSkillInterface> */
private array $skills = [];
public function register(string $class, AiSkillInterface $skill): void
{
$this->skills[$class] = $skill;
}
public function get(string $class): AiSkillInterface
{
return $this->skills[$class]
?? throw new \InvalidArgumentException("Unknown skill: $class");
}
/** @return string[] */
public function getRegisteredSkillClasses(): array
{
return array_keys($this->skills);
}
}
```
Tag all skill services with `messebau.ai_skill` and autoconfigure:
```yaml
# bundles/MessebauBundle/Resources/config/services.yaml
services:
_instanceof:
App\MessebauBundle\Service\AiSkill\AiSkillInterface:
tags: ['messebau.ai_skill']
App\MessebauBundle\Service\AiSkill\SkillRegistry:
calls:
- method: register
arguments: [!php/const App\MessebauBundle\Service\AiSkill\ContactExtractor::class, '@App\MessebauBundle\Service\AiSkill\ContactExtractor']
```
### Code: Field → Skill Routing Matrix
```php
// bundles/MessebauBundle/Service/AiSkill/FieldSkillRoutingMatrix.php
class FieldSkillRoutingMatrix
{
private const FIELD_SKILL_MAP = [
'briefingCustomerDoc' => [
ContactExtractor::class,
MultiFairDetector::class,
BudgetExtractor::class,
PainPointSummarizer::class,
ChecklistGenerator::class,
DeadlineDetector::class,
MaterialListExtractor::class,
],
'briefingBrandDoc' => [
BrandConsistencyChecker::class,
],
'briefingPlanDoc' => [
StandDataExtractor::class,
MaterialListExtractor::class,
DeadlineDetector::class,
],
'briefingAdditionalDoc' => [
DeadlineDetector::class,
],
'powerConfirmation' => [
ConfirmationParser::class,
DeadlineDetector::class,
BudgetExtractor::class,
],
'wlanConfirmation' => [
ConfirmationParser::class,
DeadlineDetector::class,
BudgetExtractor::class,
],
'cleaningConfirmation' => [
ConfirmationParser::class,
DeadlineDetector::class,
],
'riggingQuote' => [
BudgetExtractor::class,
DeadlineDetector::class,
],
'riggingDrawing' => [
StandDataExtractor::class,
],
'riggingMesseConfirmation' => [
ConfirmationParser::class,
DeadlineDetector::class,
],
'standApprovalDoc' => [
ConfirmationParser::class,
DeadlineDetector::class,
],
'standGallery' => [
// Only triggered on 2nd+ photo, handled in listener logic
],
'weeklyProtocols' => [
WeeklyProtocolSummarizer::class,
],
];
/** @return string[] */
public function getSkillsForField(string $fieldName): array
{
return self::FIELD_SKILL_MAP[$fieldName] ?? [];
}
public function fieldHasSkills(string $fieldName): bool
{
return !empty($this->getSkillsForField($fieldName));
}
}
```
### Deployment: Workers
```yaml
# docker-compose.yml (add to existing Pimcore services)
pimcore_ai_worker:
image: pimcore/php:8.2
command: >
sh -c "php bin/console messenger:consume ai_skills
--limit=25 --time-limit=300 --memory-limit=256M"
deploy:
replicas: 4 # 4 concurrent skill processors
restart: unless-stopped
environment:
MESSENGER_TRANSPORT_DSN: redis://redis:6379/messages/ai_skills
AI_SKILLS_ENABLED: 'true'
AI_LLM_PROVIDER: openai
AI_LLM_API_KEY: ${OPENAI_API_KEY}
volumes:
- ./:/var/www/html
depends_on:
- redis
```
```ini
# Or via Supervisor (for non-Docker deployments)
[program:messebau-ai-worker]
command=php /var/www/html/bin/console messenger:consume ai_skills --limit=25 --time-limit=300 --memory-limit=256M
numprocs=4
process_name=%(program_name)s_%(process_num)02d
autostart=true
autorestart=true
```
### LLM-Heavy Skills: Hybrid Approach (Tiers 2-3)
For skills requiring significant LLM calls (ConfirmationParser, BrandConsistencyChecker, LeadMagnetIdeator, AcceptanceReportGenerator, WeeklyProtocolSummarizer), PHP workers block during HTTP calls. The hybrid approach offloads LLM calls:
```
PHP Worker (messenger:consume)
│
│ Text extraction, regex preprocessing
│ Assembles LLM prompt
│ Dispatches to LLM queue
▼
LLM Transport (Redis, separate queue: ai_llm)
│
▼
Node.js/Python LLM Worker
│ Calls OpenAI/Anthropic API (async, non-blocking I/O)
│ Writes result to Redis cache key: llm:result:{messageId}
▼
PHP Worker (messenger:consume, result handler)
│ Reads from Redis
│ Calls SuggestionService::store()
```
**When to use hybrid vs. direct LLM in PHP:**
- Tier 1 skills use regex primarily, LLM only as fallback → direct PHP is fine
- Tier 2-3 skills use LLM per invocation → hybrid preferred for throughput
- Config toggle: `AI_LLM_EXECUTION_MODE = direct | hybrid` per skill or per tier
### Execution Guarantees
| Guarantee | How |
|---|---|
| At-least-once delivery | Messenger transport (Redis/Doctrine) acknowledges after handler completes |
| Failed message retry | Messenger failure transport; 3 retries with exponential backoff |
| Graceful skill failure | `try/catch` in handler; failed skill is logged, other skills continue |
| Duplicate detection | Handler checks if asset still exists + if result already stored for this asset+skill combination |
| Dead letter queue | Messages failing all retries go to `ai_skills_failed` for manual review |
| No double-processing | SuggestionService deduplicates: same asset+skill+field → updates existing instead of creating duplicate |
---
## SuggestionService API
### Data Model
```json
{
"aiSuggestions": [
{
"id": "c5f8a1b2-...",
"skill": "ContactExtractor",
"field": "customer.contactEmail",
"suggestedValue": "anna@acme.de",
"currentValue": "info@acme.de",
"confidence": 0.95,
"sourceAssetId": 12345,
"sourceAssetName": "briefing_acme_2026.pdf",
"status": "pending",
"createdAt": "2026-06-09T14:30:00Z"
}
]
}
```
### Endpoints
| Method | Endpoint | Description |
|---|---|---|
| `GET` | `/api/messebau/jobs/{id}/suggestions` | All pending suggestions for this job |
| `PATCH` | `/api/messebau/jobs/{id}/suggestions/{suggestionId}/apply` | Apply — writes suggestedValue to target field, status = `applied` |
| `PATCH` | `/api/messebau/jobs/{id}/suggestions/{suggestionId}/ignore` | Ignore — status = `ignored`, hidden from UI |
| `PATCH` | `/api/messebau/jobs/{id}/suggestions/apply-all` | Bulk apply all pending suggestions (with confirmation dialog) |
| `PATCH` | `/api/messebau/jobs/{id}/suggestions/ignore-all` | Bulk ignore |
---
## Frontend Components
### SuggestionBadge
```
┌─────────────────────────────────────────────────────────────┐
│ 💡 AI suggestion (95% confidence) [×] │
│ From: briefing_acme_2026.pdf │
│ Field: Contact Email │
│ Current: info@acme.de │
│ Suggested: anna@acme.de │
│ [Apply] [Ignore] │
└─────────────────────────────────────────────────────────────┘
```
Positioned inline below the relevant form input. Yellow left border. Non-blocking — form remains editable underneath.
### SuggestionCounter (Tab header badge)
Each tab header shows a badge if there are pending suggestions for fields within that tab:
```
┌──────────────────────────────────────┐
│ 📋 Stammdaten & Termine [2 pending] │
│ 📄 Briefing [1 pending] │
│ 🔧 Organisation │
│ 🏭 Produktion │
│ ✅ Dokumente │
└──────────────────────────────────────┘
```
---
## LLM Configuration
| Config Key | Default | Description |
|---|---|---|
| `AI_SKILLS_ENABLED` | `true` | Master toggle for all AI skills |
| `AI_LLM_PROVIDER` | `openai` | `openai` \| `anthropic` \| `ollama` (self-hosted) |
| `AI_LLM_MODEL` | `gpt-4o-mini` | Model identifier |
| `AI_LLM_API_KEY` | (env) | API key |
| `AI_LLM_BASE_URL` | `https://api.openai.com/v1` | For self-hosted (Ollama, vLLM) |
| `AI_SKILL_TIER` | `1` | Which tier of skills are active (1, 2, 3) |
| `AI_MIN_CONFIDENCE` | `0.70` | Minimum confidence to surface a suggestion |
| `AI_MAX_TOKENS_PER_CALL` | `4096` | LLM token limit per skill invocation |
When `AI_LLM_PROVIDER` is unset or unreachable, skills fall back to regex/pattern extraction only. Structured outputs (JSON) are enforced via LLM function calling or `response_format: json_object`.
---
## Event Listener
**Class**: `bundles/MessebauBundle/EventListener/AiSkillTriggerListener.php`
**Listens to**: `pimcore.asset.postAdd`
**Flow**:
```
1. Check AI_SKILLS_ENABLED config
2. DocumentClassifier → type
3. If type === 'other' → return (no skills applicable)
4. Load applicable skills for type + current AI_SKILL_TIER
5. For each skill (parallel via Symfony Messenger if async desired):
a. Extract text from asset (PDF/DOCX parser)
b. Run extraction logic (regex → LLM fallback if needed)
c. If confidence >= AI_MIN_CONFIDENCE:
- SuggestionService::add(asset, skill, field, value, confidence)
6. Dispatch message for async processing (prevent upload request blocking)
```
---
## Phased Rollout Summary
| Phase | Skills | Dependencies | Effort |
|---|---|---|---|---|
| **12a (Tier 1)** | DocumentClassifier, **DeadlineDetector (enhanced German)**, **MultiFairDetector**, ContactExtractor, **BudgetExtractor**, **PainPointSummarizer**, ChecklistGenerator | PDF text extraction, optional LLM API | 3 weeks |
| **12b (Tier 2)** | StandDataExtractor, **DocumentVersionDetector**, ConfirmationParser, BrandConsistencyChecker, DocumentMismatchDetector | DWG parsing, LLM vision model (for photos) | 3 weeks |
| **12c (Tier 3)** | MaterialListExtractor, **LeadMagnetIdeator**, AcceptanceReportGenerator, SetupProgressAnalyzer, WeeklyProtocolSummarizer | Gotenberg (already in place), LLM vision model | 3 weeks |
**New skills in bold** — discovered from analyzing real trade fair documents in `plan/documents/`: