Profile
Back to NewsBack
Dev.to 9 min
Reader Mode
WildProof: Go Outside With a Question, Come Back With Evidence

WildProof: Go Outside With a Question, Come Back With Evidence

1 day ago

FieldProof implementation plan

The workspace is currently empty, so this should be treated as a greenfield MVP. The deadline is approximately 28.5 hours away: Oct 11, 2026 at 11:59 PM PDT, which is Oct 12 at about 12:29 PM IST.

The priority is not building every feature. It is producing a polished, credible demo and a strong DEV write-up.


1. Product definition

Project

FieldProof

One-sentence pitch

FieldProof is an offline-first open-AI app that turns a short walk into structured, uncertainty-aware field notes.

Core user flow

Choose a mission
    ↓
Walk outside
    ↓
Capture photos + notes
    ↓
Run local/open AI analysis
    ↓
Review evidence and confidence
    ↓
Export/share a field report

MVP missions

Implement only three:

  1. Habitat comparison

    • Capture one sunny-area observation and one shaded-area observation.
    • Generate a comparison report.
  2. Pattern hunt

    • Find three visually different natural objects or textures.
    • AI groups them into broad categories without claiming exact species.
  3. Sound and scene

    • Capture a short audio clip or text description.
    • Generate a structured environmental note.

Avoid building a general-purpose nature identification app.


2. Recommended stack

Use a stack optimized for speed and demo quality:

Layer Choice
Frontend React + Vite + TypeScript
Styling Tailwind CSS
Local persistence IndexedDB via idb
AI Gemma through Ollama or a hosted open-weight endpoint
Backend Small Python FastAPI service
Tabular model TabPFN
Durable workflow Temporal, only for the analysis pipeline
Observability Sentry
Deployment Render
Storage Local browser storage for MVP; MongoDB/Tiger Data only if time remains
Testing Vitest + Playwright smoke test

If the environment already has a preferred stack installed, preserve it. Otherwise, React/Vite/TypeScript is the fastest reliable route.


3. Architecture

React PWA
 ├─ Mission selection
 ├─ Camera/file/audio capture
 ├─ Evidence review
 ├─ Field report
 └─ Offline queue
        │
        ▼
FastAPI API
 ├─ /missions
 ├─ /analyze
 ├─ /predict-quality
 └─ /health
        │
        ├─ Gemma inference
        ├─ TabPFN quality prediction
        ├─ Temporal workflow
        └─ Sentry tracing

Important architecture decision

The app should still function without the backend:

  • Mission templates are bundled locally.
  • Evidence is saved in IndexedDB.
  • The user can complete a walk offline.
  • AI analysis runs when the backend is available.
  • Failed analysis is queued for retry.

This gives FieldProof a genuine offline-first story rather than merely adding “offline” to the README.


4. Data model

Start with simple typed objects.

type MissionType =
  | "habitat-comparison"
  | "pattern-hunt"
  | "sound-and-scene";

type EvidenceKind = "photo" | "audio" | "note";

type Confidence = "observed" | "inferred" | "unknown";

interface EvidenceItem {
  id: string;
  kind: EvidenceKind;
  localUrl?: string;
  text?: string;
  capturedAt: string;
  latitude?: number;
  longitude?: number;
}

interface Expedition {
  id: string;
  missionType: MissionType;
  startedAt: string;
  completedAt?: string;
  evidence: EvidenceItem[];
  status: "draft" | "queued" | "processing" | "complete" | "failed";
}

interface FieldReport {
  summary: string;
  observations: Array<{
    claim: string;
    confidence: Confidence;
    supportingEvidenceIds: string[];
  }>;
  missingEvidence: string[];
  suggestedNextStep: string;
  qualityScore?: number;
}

The observed / inferred / unknown distinction is central to the product and should appear visibly in the UI.


5. Build sequence

Phase 1 — Project foundation

Time box: 45 minutes

Create:

  • Vite React TypeScript app.
  • Tailwind setup.
  • Basic routing or screen-state navigation.
  • ESLint and formatter.
  • .env.example.
  • MIT license.
  • README skeleton.

Initial screens:

Home
Mission selection
Active expedition
Evidence review
Field report
About / methodology

Do not spend time on authentication.


Phase 2 — Mission and evidence capture

Time box: 3 hours

Implement:

  • Mission cards.
  • Start mission button.
  • Photo upload using file input.
  • Optional camera capture on mobile.
  • Text observation input.
  • Optional audio upload.
  • Evidence list with delete/reorder.
  • Save draft to IndexedDB.
  • Resume unfinished expedition after refresh.

Acceptance criteria:

  • A user can start a mission without an account.
  • Refreshing the page does not lose the expedition.
  • A user can complete a mission with photos and notes even while offline.

Phase 3 — Report generation without AI

Time box: 1.5 hours

Before connecting a model, build a deterministic mock report generator.

This allows the complete UI flow to work immediately:

Evidence review
    ↓
Generate report
    ↓
Report page with confidence labels
    ↓
Export JSON

This is important because the AI integration may be the part most likely to fail under deadline pressure.

The fallback must be explicit in the UI:

“AI analysis unavailable. Showing a locally generated evidence summary.”

Do not silently present mock output as AI output.


Phase 4 — Gemma integration

Time box: 3 hours

Create a small FastAPI service with:

GET  /health
POST /analyze
POST /predict-quality

POST /analyze should receive:

  • Mission type.
  • Text observations.
  • Image references or compressed images.
  • Optional audio transcript.
  • Location metadata only if the user opts in.

Prompt Gemma to return strict JSON:

{
  "summary": "...",
  "observations": [
    {
      "claim": "...",
      "confidence": "observed",
      "evidence_ids": ["..."]
    }
  ],
  "missing_evidence": ["..."],
  "suggested_next_step": "..."
}

Add validation with Pydantic. If the model returns invalid JSON:

  1. Log the failure.
  2. Retry once with a repair prompt.
  3. If it still fails, return an explicit analysis error.
  4. Preserve the evidence so the user can retry.

Do not let malformed AI output break the whole expedition.


Phase 5 — TabPFN category implementation

Time box: 2 hours

Use TabPFN for a narrow, measurable task:

Predict whether an expedition contains enough evidence for a useful field report.

Create a small structured feature set:

  • Number of photos.
  • Number of notes.
  • Note length.
  • Number of distinct mission locations.
  • Audio duration.
  • Time spent on mission.
  • Lighting/habitat metadata.
  • Evidence completeness.

Create a baseline:

quality = evidence_count >= 3 && note_length >= 40

Compare the baseline against TabPFN using a small labeled fixture dataset.

The article should show:

  • Dataset size.
  • Feature list.
  • Baseline performance.
  • TabPFN performance.
  • One limitation.

If real model integration becomes unstable, keep the TabPFN experiment as a standalone reproducible script and do not fake live predictions in the UI.


Phase 6 — Temporal workflow

Time box: 2 hours

Use Temporal only around the analysis pipeline:

analyzeExpedition
  ├─ normalizeEvidence
  ├─ runGemmaAnalysis
  ├─ validateReport
  ├─ runTabPFNQualityPrediction
  └─ persistReport

Demonstrate one failure scenario:

  1. Start processing.
  2. Force one activity to fail.
  3. Show retry/resumption.
  4. Complete the report without restarting the expedition.

This gives you a strong Best Use of Temporal story.

If Temporal setup threatens the deadline, retain the workflow interface and document the durable pipeline separately rather than delaying the entire app.


Phase 7 — Sentry instrumentation

Time box: 45 minutes

Instrument:

  • Expedition creation.
  • Evidence upload.
  • Gemma analysis duration.
  • Model parse failures.
  • TabPFN prediction duration.
  • Workflow retries.

Capture screenshots showing:

  • A successful analysis trace.
  • A failed/retried analysis.
  • Latency or error information.

This supports Best Use of Sentry Agent Tracing.

Never include API keys or private location data in traces.


Phase 8 — UI polish and demo path

Time box: 3 hours

Polish only the main path:

  1. Home.
  2. Mission selection.
  3. Evidence capture.
  4. Report.
  5. Export.

Add:

  • Strong empty states.
  • Loading state with clear progress.
  • Offline indicator.
  • Retry button.
  • Confidence badges.
  • “Why this claim?” evidence links.
  • Mobile layout.
  • One sample expedition for judges.

The report screen is the most important screen. Make the AI’s uncertainty visually obvious.


6. Deployment plan

Render deployment

Deploy:

  • Frontend as a static site.
  • FastAPI as a web service.
  • Model service separately only if required.

Required production checks:

/health returns 200
frontend can call production API
analysis errors display correctly
no localhost URLs remain
README setup instructions work

If local Gemma cannot run reliably on Render, use:

  • Local Gemma for the development/demo evidence.
  • A clearly documented open-weight inference endpoint for the hosted demo.
  • Or a deterministic demo fixture with an explicit “demo mode” label.

Do not imply that a hosted fallback is local inference if it is not.


7. Testing checklist

Automated tests

Minimum:

  • Mission creation.
  • IndexedDB save/resume.
  • Report schema validation.
  • Invalid model JSON handling.
  • TabPFN feature generation.
  • Offline queue behavior.
  • API health endpoint.

Manual smoke test

Run this exact script:

  1. Open deployed app on mobile viewport.
  2. Start Pattern Hunt.
  3. Upload three images.
  4. Add a 50-word note.
  5. Disconnect network.
  6. Refresh page.
  7. Confirm evidence remains.
  8. Reconnect network.
  9. Run analysis.
  10. Confirm every report claim links to evidence.
  11. Export JSON.
  12. Confirm no secrets appear in output.

8. Prize category evidence

Add a section to the README called Prize Category Evidence.

Category Evidence to include
Overall Live demo, screenshots, architecture, polished write-up
Gemma Model name, prompt, local/open inference explanation, sample output
TabPFN Dataset, features, baseline comparison, result chart
Temporal Workflow diagram and retry/resume recording
Sentry Trace screenshots and performance findings
Render Production URL and deployment architecture
Entire Link/embed to the actual development session
ElevenLabs Only if narration is a meaningful feature
MongoDB Atlas Only if historical observation search is implemented

Do not enter a category just because a library was imported. Every category needs a visible feature and an explanation in the article.


9. Submission article plan

The article is crucial because the challenge explicitly weights writing quality heavily.

Suggested title

I Built an Offline AI That Turns a 20-Minute Walk Into Evidence

Article structure

  1. The problem

    • Most outdoor apps give reminders or routes.
    • They do not help users notice and record their environment.
  2. The product

    • Explain one expedition from start to finish.
  3. Why open AI

    • Privacy.
    • Offline operation.
    • Inspectable behavior.
    • Lower dependence on proprietary APIs.
  4. Architecture

    • Include the system diagram.
  5. Gemma

    • Show structured output and uncertainty handling.
  6. TabPFN

    • Show the prediction task and baseline comparison.
  7. Temporal

    • Show failure recovery.
  8. Sentry

    • Show what you measured and what you fixed.
  9. Limitations

    • Small dataset.
    • No authoritative species claims.
    • Model performance varies by image/audio quality.
  10. How to run it

    • Repository.
    • Environment variables.
    • Local model setup.
    • Demo URL.
  11. Prize categories

    • List only categories genuinely supported.

10. Time budget from now

Time remaining Deliverable
0–4 hours Foundation, mission flow, IndexedDB
4–8 hours Evidence capture and complete mock report flow
8–12 hours Gemma API and validated reports
12–14 hours TabPFN experiment
14–16 hours Temporal retry demo
16–17 hours Sentry instrumentation
17–20 hours UI polish and deployment
20–24 hours Tests, screenshots, article
Final 4 hours Buffer, submission checks, publish before cutoff

If behind schedule, cut in this order:

  1. ElevenLabs.
  2. MongoDB/Tiger Data.
  3. Map/history features.
  4. Audio capture.
  5. Complex Temporal deployment.

Do not cut:

  • Evidence-linked reports.
  • Confidence labels.
  • Gemma integration.
  • TabPFN comparison.
  • Deployment.
  • Article quality.

Definition of done

FieldProof is ready when:

  • A new user can complete one mission in under five minutes.
  • Evidence survives refresh and temporary disconnection.
  • The report distinguishes observed, inferred, and unknown claims.
  • Gemma is used for the central AI task.
  • TabPFN has a real baseline comparison.
  • At least one Temporal retry is demonstrable.
  • Sentry contains useful traces.
  • The app is deployed.
  • The repository has a license and reproducible setup.
  • The DEV article includes screenshots, architecture, results, limitations, and category evidence.
  • The submission is published before Oct 11, 2026 at 11:59 PM PDT.

Community Wisdom: I Taught Local AI to Help Me Notice the World

The implementation should not compete as another generic “AI helps you notice nature” app. The differentiation needs to be the structured evidence workflow, explicit uncertainty, and measurable analysis quality.

Community Wisdom: EcoID: An Offline Plant Identifier That Gets You Outside

An offline plant-identification angle is already represented. FieldProof should instead remain useful when identification is uncertain, making evidence collection and comparison the core product rather than exact recognition.

Chat with me