# TASKS_PROMPTS.md - Viva Mais AI implementation spec A from-scratch build, ordered by dependency. Each task below is a self-contained prompt: paste it to a coding agent verbatim. Always prepend the line: > Follow AGENTS.md. Work test-first (TDD for units, Gherkin for use cases). > Make one or more atomic commits. Keep the suite green. Use cases referenced as "UC#" map to USE_CASES.md. Build order is not value order: foundations and extraction come first, intelligence and assembly after. Tech baseline: Python 3.13, uv (env + tool runner), ruff (lint + format), mypy (strict types on src), Gradio, pytest + pytest-bdd, SQLite, llama.cpp (MiniCPM-V 4.6 GGUF), whisper, Modal (dev bench only). Clean Architecture layers: `domain` -> `application` (interactors + ports) -> `adapters` -> frameworks. The deployed Space installs from `requirements.txt` (pip); uv and pyproject are for local development and CI only. --- ## Phase 0 - Foundation ### T0.1 Repository scaffold and tooling Objective: create the project skeleton and test harness so every later task can run test-first. Do: 1. Add `pyproject.toml` (src layout, package `vivamais` under `src/`): runtime deps `gradio`, `numpy`, `pillow`; a `models` extra (`llama-cpp-python`, `huggingface_hub`); a uv `dev` dependency-group (`pytest`, `pytest-bdd`, `ruff`, `mypy`); pytest configured with `pythonpath = ["src"]` and testpaths `tests/unit`, `tests/steps`; ruff and strict-mypy config. Add a `requirements.txt` with the Space runtime deps (the Space installs it via pip; keep it synced with `uv export --no-dev`). 2. Create the layer directories with `__init__.py`: `src/vivamais/domain`, `src/vivamais/application`, `src/vivamais/adapters/{models,persistence,ingest,presenters}`, plus `features/`, `tests/unit/`, `tests/steps/`. 3. Add `Makefile` targets driving uv: `install` (uv sync), `test` (uv run pytest), `lint` (ruff check + format check), `types` (mypy), `check` (all three), `run` (Space with the mock backend), `lock-requirements`. 4. Add `.gitignore` covering privacy paths: `raw_data/`, `*_chat.txt`, `*.opus`, `*.vcf`, `data/`, `outputs/`, `*.db`, `__pycache__/`. 5. CI (`make check` is the real gate). Hugging Face does not run GitHub Actions, so add a workflow only if the repo is mirrored to GitHub for development; it would run uv sync, ruff, mypy, and pytest, and could auto-deploy to the Space. 6. Add one trivial passing unit test to prove the harness and the src import path. Done when: `make check` (lint, types, tests) passes clean. ### T0.2 Domain entities and value objects Objective: the framework-free core model. Do: 1. Implement a `Money` value object: integer cents plus currency; a parser that handles Brazilian formats ("R$ 2.350,00", "2.350" meaning 2350, "2,35" meaning 2.35, plain "50", and numeric input) into cents; a formatter (cents -> "R$2.350,00"). Cover the thousands-vs-decimal ambiguity with tests. 2. Implement frozen dataclasses: `Customer` (name, contacts, fellow_travelers, facts, trips, last_contact), `Trip` (locator, passengers, airline, total: Money, segments), `Flight` (airline, number, origin, destination, depart, arrive, cabin, seat), `Payment` (payer, amount: Money, date, method, bank), `Message` (contact, kind, text, media_path, timestamp, from_me), `Document` (passenger, kind, status), `Report`/`Issue` (validation result). 3. Implement enums: `PipelineStage` (QUOTE, DOCUMENTS, PAYMENT, TICKETING, TRAVEL, DONE) and `NextActionKind`. Constraints: no imports outside the standard library and `domain`. Add a test that asserts the domain package imports nothing from `application`, `adapters`, or any third-party framework. Done when: domain unit tests green (Money parse/format, entity equality). ### T0.3 Application ports Objective: the abstract boundaries the interactors depend on. Do: in `application/ports.py` define `typing.Protocol` interfaces: - `VisionModel.extract(image, schema_hint: str) -> dict` (image -> structured fields). - `TextModel.generate(messages: list[dict]) -> str`. - `Transcriber.transcribe(audio_path: str) -> str`. - `Repository` with save/get/list for customers, payments, trips, documents, messages, plus query methods used by UC8. - `Clock.now() -> str` (ISO timestamp). Also define the message content-part shape (text, image, audio) used by `TextModel`. Done when: protocols compile and a throwaway dummy satisfies each shape. ### T0.4 Mock adapters and fixtures Objective: deterministic adapters so all later tests run offline. Do: 1. `adapters/models/mock.py`: `MockVisionModel` returns canned, schema-valid receipt/ticket JSON selected by `schema_hint`; `MockTextModel` returns a canned reply; `MockTranscriber` returns canned Portuguese text. 2. A `FixedClock`. 3. `tests/conftest.py` exposing these as fixtures plus an in-memory `Repository`. Done when: a smoke test exercises each mock through its port. ### T0.5 WhatsApp export ingestion Objective: turn a raw export into domain `Message` objects. Do: `adapters/ingest/whatsapp_export.py`: 1. `load_zip(path) -> dir` unzips an export to a temp directory. 2. `parse_chat(text, owner) -> list[Message]`: parse iPhone bracketed lines `[DD/MM/YYYY, HH:MM:SS] Sender: text` and the Android dash variant; join continuation lines into the previous message; classify media from `` / `` (kind by extension) and from omitted markers; skip the end-to-end-encryption system line; set `from_me` when the sender equals `owner`. 3. `counterpart(messages, owner)` returns the most frequent non-owner sender. Test-first with fixture snippets covering: iPhone, Android, attached image, attached pdf with extra label text, omitted audio, a multi-line message, and the system line. Done when: parser unit tests green. ### T0.6 SQLite repository Objective: a concrete `Repository` over SQLite. Do: `adapters/persistence/sqlite_repository.py` implementing the `Repository` port; tables for customers, messages, payments, trips, documents, pipeline; round-trip nested data (trip segments, customer facts). Test save/get/list and nested round-trips. Done when: repository tests green and it satisfies the port. ### T0.7 Gradio shell and composition root Objective: wire adapters to interactors and render an empty dashboard. Do: `app.py` at the repo root (matching the Space `app_file: app.py`; there is no separate `space/` directory, the repo root is the Space): 1. `build_container(env)` factory selecting adapters (`mock` default; `llamacpp`, `whisper`, `modal` later) and constructing the interactors. 2. A Gradio `Blocks` UI (Portuguese strings) with an upload control and empty dashboard tabs: Cliente, Pagamentos, Viagens, Funil, Documentos, Linha do tempo, Busca. No use-case logic yet. Done when: the app builds warning-free, the factory returns a mock-backed container, and a UI smoke test passes. --- ## Phase 1 - Extraction use cases ### T1 Voice transcription and timeline (UC6) Write first: `features/timeline.feature` ```gherkin Scenario: Build a searchable timeline from a thread with voice notes Given a parsed conversation that includes two voice notes When the user opens the conversation Then every voice note is transcribed And the timeline lists all events in chronological order ``` Build: `application/transcribe_timeline.py` with `BuildTimeline(messages, Transcriber) -> Timeline`. Audio messages get a transcription; text messages pass through; other media become labeled events. Add a presenter mapping the timeline to dashboard rows. Sub-tasks: (1) event ordering, (2) audio -> transcription via the port, (3) a short per-conversation summary line. Done when: the scenario and unit tests are green on mocks. ### T2 Payment tracking and reconciliation (UC2) Write first: `features/payments.feature` ```gherkin Scenario: Show paid, outstanding, and profit for a customer Given a customer thread with a quote of "R$ 3.480,00" and a receipt of "R$ 2.350,00" When the user opens the customer card Then the total paid is "R$2.350,00" And the outstanding balance is "R$1.130,00" ``` Build: `application/track_payments.py`: 1. `ExtractPayments(images, VisionModel) -> list[Payment]` (parse PIX/TED receipts into `Payment` with `Money`). 2. `Reconcile(payments, sale_total, cost) -> Ledger` with `total_paid`, `outstanding = sale_total - total_paid`, and `profit = sale_total - cost` (cost optional; mark unknown when absent), matching a receipt's payer to a trip passenger. Test the `Money` arithmetic, the payer-to-trip match, and the outstanding/profit calculations. Done when: the scenario and units are green. ### T3 Trip extraction, summary, and reminders (UC4) Write first: `features/trip_summary.feature` (Given an e-ticket with locator, passengers, and flights; When processed; Then a structured trip, a summary message, check-in reminders, and an .ics file are produced). Build: `application/extract_trip.py`: 1. `ExtractTrip(image_or_pdf, VisionModel) -> Trip` (validate required fields, normalize IATA codes and datetimes). 2. `GenerateTripSummary(trip, persona, TextModel) -> str` in her template. 3. `BuildReminders(trip, Clock) -> list[Reminder]` (check-in windows). 4. `trip_to_ics(trip) -> str` (one VEVENT per segment). Test trip validation, the .ics event count and timestamp format, and that the summary uses the persona template. Done when: the scenario and units are green. ### T4 Document-collection tracking (UC7) Write first: `features/documents.feature` (Given a trip whose passengers must each submit RG and CPF and a thread where one passenger's documents arrived; When the user prepares ticketing; Then the system shows which are received and which are missing). Build: `application/track_documents.py`: detect document submissions from the thread (VisionModel classifies RG/CPF document images, plus text mentions), cross-reference with the trip's passengers, and produce a per-passenger received/missing list. Test received-vs-missing given a passenger list and a set of detected documents. Done when: the scenario and units are green. --- ## Phase 2 - Assembly and intelligence ### T5 Auto-built customer CRM card (UC1, the centerpiece) Write first: `features/crm_card.feature` ```gherkin Scenario: Assemble a customer card from a raw export Given a raw WhatsApp export with messages, a receipt image, a ticket image, and a voice note When the user uploads it to the Space Then a customer card is created with the customer's identity and contacts And it lists the trips discussed And it includes the payment ledger and the conversation timeline ``` Build: `application/build_crm_card.py` with `BuildCrmCard` orchestrating T1-T4 plus identity extraction (customer name, contacts, fellow travelers) into a `Customer` aggregate, persisted via `Repository`. Wire it to the app upload flow and render the Cliente card via a presenter. Sub-tasks: (1) identity and fellow-traveler extraction, (2) link trips and payments to the customer, (3) attach the timeline, (4) persist and render. Done when: the upload-to-card scenario is green on mocks and the card renders. ### T6 Sales-pipeline status and next action (UC3) Write first: `features/pipeline.feature` ```gherkin Scenario: Surface the next action when documents are missing Given a customer who has paid in full but is missing one passenger's RG When the user opens the customer Then the pipeline stage is "DOCUMENTS" And the next action is to collect the missing document ``` Build: `application/pipeline_status.py` with `InferPipeline(customer, Clock) -> (stage, remaining_steps, next_action)`. Deterministic rules: missing documents -> collect documents; outstanding balance -> request payment; ticket issued and trip upcoming -> send check-in reminder; quote sent with no reply for N days -> follow up. Use `Clock` for the day-based rules. Test each rule and the stage inference from aggregate state. Done when: the scenario and unit tests are green. ### T8 Cross-customer dashboard and search (UC7) Write first: `features/dashboard.feature` (Given several processed customers; When the user opens the overview; Then aggregated money, time, and profit are shown and a question like "who still owes me?" returns the right customers). Build: `application/dashboard.py`: `Aggregate(repo) -> totals` (money in, outstanding, profit, time per customer) and `Search(query, repo) -> answer` using structured queries, with an optional `TextModel` step to map a natural question to a filter. Render the Busca and overview panels. Test the aggregation sums and two canned queries ("who owes", "trips this week"). Done when: the scenario and units are green. --- ## Phase 3 - Real model adapters ### T9 llama.cpp vision and text adapter (MiniCPM-V 4.6) Objective: implement `VisionModel` and `TextModel` against MiniCPM-V 4.6 GGUF. Do: `adapters/models/llama_cpp.py`. Resolve the language GGUF (`MiniCPM-V-4.6-Q4_K_M.gguf`) and the projector (`mmproj-MiniCPM-V-4.6-F16.gguf`) from `openbmb/MiniCPM-V-4.6-gguf`; run via `llama-cpp-python` (or a `llama-server` subprocess) with the mtmd/vision path; image + JSON-mode prompt -> structured text; text-only prompts for replies. Requires a llama.cpp build at release `b9049` or later. Provide a guarded integration test (skipped without the model) and a contract test proving parity with the mock's interface. Done when: the adapter satisfies both ports and the guarded test runs locally. ### T10 whisper transcriber adapter Objective: implement `Transcriber` for Portuguese voice notes. Do: `adapters/models/whisper.py` using whisper.cpp or faster-whisper; handle `.opus`/`.m4a`; Portuguese; return plain text. Guarded integration test plus mock parity. Done when: the adapter satisfies the port and the guarded test runs. ### T11 Modal dev bench (llama.cpp serve) and client Objective: run the model on Modal for curation and evals only. Do: `modal/serve.py` running llama.cpp (llama-server) over the MiniCPM GGUF as an OpenAI-compatible, scale-to-zero, secret-authenticated endpoint; `adapters/models/modal_client.py` implementing the ports by calling it. This is the dev bench, never a dependency of the deployed Space. Done when: the script compiles and the client adapter passes the contract test. --- ## Phase 4 - Vision fine-tuning and evals ### T12 Vision extraction fine-tune (MiniCPM-V 4.6) Objective: stop OCR hallucinations on Brazilian PIX receipts, airline tickets, and RG/CPF images. Publish a vision LoRA or merged GGUF the Space can load (Well-Tuned badge). Modal trains; llama.cpp serves in the Space. **Step 1 — Gold labels (local, never committed)** - Export WhatsApp zips stay in `raw_data/` (gitignored). - For each relevant image, hand-verify JSON matching `ReceiptExtraction`, `TicketExtraction`, and `DocumentExtraction` schemas (same fields as `adapters/models/prompts.py`). - Tag negatives: chat screenshots, memes, and non-receipt images should map to `type=unknown` or empty actionable fields. - Target 50–200 labeled images before training; prioritize failure cases from production traces (BRAZILIANTRAVEL*, placeholder passengers, wrong amounts). **Step 2 — Redact before any publish** - `finetune/redact.py`: strip CPF, phone, email, card numbers, client names, and blur or drop RG/CPF document pixels before images enter a Hub dataset. - Reuse rules from `adapters/observability/redact.py`; extend for image-side redaction. Never commit raw exports or identity documents. **Step 3 — Build vision dataset** - `finetune/build_dataset.py`: emit LLaMA-Factory multimodal JSON (image path + user prompt + assistant JSON). Use the same prompts as production (`COMBINED_PROMPT`, `TICKET_PROMPT`, `RECEIPT_PROMPT`) so train/serve match. - Include hard negatives and multi-field combined examples (receipt-only, ticket-only, unknown). - Split train/val (e.g. 90/10); register a private Hub dataset revision. **Step 4 — Baseline eval (before training)** - `evals/run_evals.py` on the base `MiniCPM-V-4_6-Q4_K_M.gguf` with gold cases. - Record per-field accuracy, locator F1, false-trip rate (hallucinated rows that pass `is_actionable_ticket`), and amount exact-match. **Step 5 — Train on Modal** - `finetune/train.py` + `modal/train.py`: QLoRA (or full SFT if small enough) over `openbmb/MiniCPM-V-4.6` with LLaMA-Factory; vision projector frozen or lightly tuned per OpenBMB recipe. - Hyperparams starting point: 2–4 epochs, lr 1e-4, rank 64, batch size tuned to GPU VRAM. Dry-run mode in CI without GPU. **Step 6 — Merge, GGUF, deploy** - `finetune/merge_export.py`: merge LoRA into base, convert to GGUF (Q4_K_M), upload to `marinarosa/vivamais-vision` (or similar) on the Hub. - Point the Space `VARIANT.repo` / env override at the tuned GGUF; keep mmproj unchanged unless OpenBMB docs say otherwise. **Step 7 — Regress and ship** - Re-run evals; tuned model must beat base on locator F1 and cut false trips. - Spot-check Esther export on Space before marking done. Test redaction, dataset schema validation, and train dry-run. Done when: fixture dataset builds, train dry-runs, and eval harness shows base-vs-tuned comparison. ### T13 Evals harness Objective: prove extraction quality and track fine-tune progress. Do: `evals/metrics.py` (per-field extraction accuracy, ticket locator/passenger/ segment F1, false-actionable rate, CER for transcription, reconciliation precision/recall/F1); `evals/run_evals.py` running each task over gold cases through a chosen adapter, with base-vs-tuned comparison; `evals/compare.py` leaderboard. Bundle tiny gold cases aligned with the mock outputs so it runs green offline. Test each metric directly and `run_evals` on the mock adapter. Done when: metric units are green and the harness runs end to end on mocks. --- ## Phase 5 - Polish and submission ### T14 Off-Brand UI and dashboard polish Objective: an elegant, mobile-first CRM dashboard (Off-Brand badge). Do: a custom theme plus CSS or a `gr.Server` custom frontend; assemble all panels (card, payments, trips, pipeline with next action, documents, timeline, search) cleanly; add loading states for model calls; show the per-step tracer in a panel. Keep all view logic in presenters; the UI holds no business rules. Done when: the app builds, the dashboard renders a full customer end to end on mocks, and the custom frontend is in place. ### T15 Deployment and submission Objective: ship it. Do: fill the Space card (emoji, color gradient, `short_description`, Backyard AI tag); deploy under the hackathon org; export agent traces to the Hub (Sharing is Caring); record the 60 to 90 second demo video (outcome-first hook, real footage, captions); publish the social post tagging Gradio, Hugging Face, and OpenBMB; write the Field Notes blog post. Complete the submission checklist (Space link, video, social post). Done when: the Space is live and the submission checklist is complete. --- ## Suggested order and parallelism T0.1 -> T0.2 -> (T0.3, T0.4) -> (T0.5, T0.6) -> T0.7 unblock everything. Phase 1 (T1-T4) can run in parallel once Phase 0 is done; T5 depends on T1-T4. T6-T8 depend on T5. Phase 3 adapters can be built in parallel with Phases 1-2 because the interactors only depend on the ports. T12-T13 need real data and the adapters. Phase 5 is last.