Visiolog Platform Documentation

Engineering reference, API specifications, and operational manuals for the full-stack Visiolog document intelligence and 2D spreadsheet platform.

Architecture Notice: Demo vs Production Application

The browser demonstration hosted on this domain is an intentional, zero-backend showcase designed to mirror the visual ergonomics, interaction flow, and spreadsheet editing mechanics of the full platform while operating 100% client-side with bring-your-own-key settings. The complete production Visiolog platform detailed below includes persistent Supabase PostgreSQL database storage, multi-page batch document pipelines, automated schema validation rule enforcement, multi-tenant project hierarchies, and cloud vector storage.

1. System Architecture & Core Stack

Full Stack

Visiolog is architected as an end-to-end vision extraction and spreadsheet engineering platform. Built with modern TypeScript and the Next.js App Router, the application combines client-side Web Worker acceleration with serverless inference pipelines and PostgreSQL data management.

Layer Technology Functionality
Framework Next.js 15+ (App Router) Server Components, Edge API routes, streaming SSR.
Language TypeScript 5.x (Strict) End-to-end type safety, RPC contracts, Zod schemas.
Database Supabase (PostgreSQL 15+) Multi-tenant data storage, Row Level Security (RLS), real-time sync.
Inference Gemini 2.5, OpenRouter, Ollama Multi-model vision OCR, structured tabular parsing.
Spreadsheet Custom Virtualized Grid & SheetJS High-performance 2D cell editing, native Excel (.xlsx) generation.
Mobile Progressive Web App (PWA) Installable mobile client, touch gestures, offline caching.

2. Authentication & Multi-Tenant Workspaces

Supabase Auth

The production platform utilizes Supabase Authentication supporting passwordless Magic Links, OAuth providers (GitHub, Google), and secure session tokens with JSON Web Token (JWT) verification on every API route.

Workspace Hierarchy

  • Organizations: Top-level tenant container managing billing, API key limits, and member seats.
  • Projects: Logical folder groups categorizing related spreadsheets, extraction templates, and audit logs.
  • Sheets: Individual tabular documents with versioned revision histories and cell metadata.
TypeScript — Workspace Service Contract
interface WorkspaceProject { id: string; organizationId: string; name: string; documentCount: number; createdAt: string; updatedAt: string; schemaRules: ColumnValidationRule[]; }

3. Multi-Page Document Ingestion Pipeline

Client & Server Processing

Visiolog accepts paper receipts, invoices, warehouse logbooks, financial tables, and multi-page PDFs. The ingestion pipeline executes pre-flight image preprocessing to maximize OCR accuracy while minimizing network payload size:

  1. Downsampling & Compression: HTML5 Canvas downsamples high-resolution smartphone captures (12MP+) to optimal inference dimensions (2048px maximum bound) at 0.85 JPEG quality, reducing payload size by over 80%.
  2. Auto-Orientation & De-Skew: EXIF metadata extraction and edge analysis correct rotated captures automatically before vision processing.
  3. Multi-File Queueing: Batch ingestion queues allow users to load sequential pages and merge extracted records into a unified tabular sheet.

4. AI Vision Inference Router

Multi-Model Support

Visiolog features a provider-agnostic inference engine that dynamically routes extraction tasks across cloud frontier models and local self-hosted workers:

  • Google Gemini (Default): Gemini 2.5 Flash and Gemini 2.5 Pro via REST with automatic exponential backoff retry on HTTP 429 rate limits.
  • OpenRouter: Unified access to Meta Llama 3.2 Vision, Qwen 2.5 VL, and Anthropic Claude 3.5 Sonnet.
  • Local Ollama: Air-gapped on-premise execution with local models (`llama3.2-vision`, `minicpm-v`) at zero marginal API cost.
  • Custom OpenAI Endpoints: Integration with proprietary internal inference servers via standard `chat/completions` API contracts.
JSON — System Vision Prompt Payload
{ "prompt": "Extract all tabular data from this document into strict RFC 4180 CSV format. Output ONLY raw CSV starting with column headers. Do not wrap in markdown fences.", "temperature": 0.1, "maxTokens": 4096 }

5. Header Reconciliation & Schema Rules

Data Quality Engine

Raw OCR outputs often contain misspelled or inconsistently named column headers (e.g. "Qty", "Quantity", "QTY", "Amount", "Total Price"). Visiolog's reconciliation engine standardizes these fields into structured schemas:

  • Semantic Field Matching: Automatically aliases synonyms to canonical database column names.
  • Type Enforcement: Formats currency, integers, floating points, dates (ISO 8601), and SKU patterns.
  • Validation Rules: Rejects or flags malformed values using user-defined regular expressions and mandatory constraint checks.

6. 2D Interactive Spreadsheet Studio

Virtual Grid Engine

The spreadsheet studio delivers instant, low-latency table manipulation directly inside the web browser:

  • Inline Cell & Header Editing: Click any cell or column header to modify values directly with keyboard arrow navigation.
  • Column Sorting: Click column headers or the Sort tool to toggle alphanumeric sorting ascending and descending.
  • Real-Time Search Filter: Instant row filtering isolates target records without altering underlying data structures.
  • Notes View Synchronizer: Switch between 2D table grid view and structured markdown notes view seamlessly.

7. Mobile PWA & Touch Gesture Suite

Progressive Web App

On mobile viewports (< 960px), Visiolog shifts to a native-feeling mobile experience:

  • Tabbed Workspace Switcher: Direct switching between Upload, Table, Notes, and History tabs eliminates vertical scroll fatigue.
  • Monochrome Icon Controls: Space-efficient touch buttons adhere to minimum 44px hit targets.
  • Service Worker Caching: Service worker (`sw.js`) pre-caches core assets, enabling instant load times and offline review capabilities.

8. Enterprise Security & Zero-Retention Privacy

Compliance Standards

Visiolog is architected with strict enterprise data privacy safeguards:

  • Row Level Security (RLS): Every PostgreSQL table enforces tenant isolation policies ensuring users can only read and write their own documents.
  • Zero-Retention Inference Mode: Documents sent to vision providers are processed ephemerally without training or persistent storage on provider infrastructure.
  • Type-to-Confirm Data Wipe: Users and administrators can purge all local and cloud data by explicitly entering the confirmation phrase `delete my data`.

9. REST API Reference & Webhook Contracts

Integration Guide

Visiolog exposes clean RESTful endpoints for programmatic document processing:

HTTP POST /api/vision/extract
curl -X POST https://api.visiolog.com/v1/extract \ -H "Authorization: Bearer VISIOLOG_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "image": "data:image/jpeg;base64,...", "provider": "gemini", "model": "gemini-2.5-flash", "outputFormat": "json" }'
JSON Response Schema (200 OK)
{ "status": "success", "documentId": "doc_99218274", "headers": ["Item", "Quantity", "Price", "Total"], "rows": [ ["Thermal Paper", "50", "1.20", "60.00"], ["Document Scanner", "2", "450.00", "900.00"] ], "stats": { "rowCount": 2, "colCount": 4, "latencyMs": 1420 } }

10. Multi-Format Client-Side Export Engine

Data Portability

Extracted spreadsheets can be exported directly from the browser without server round-trips:

  • Microsoft Excel (`.xlsx`): Full binary workbook generation via client-side SheetJS with automatic column width formatting.
  • RFC 4180 CSV (`.csv`): Strict RFC 4180 quotation escaping for universal compatibility with databases and spreadsheets.
  • JSON Records (`.json`): Key-value paired structured objects ideal for database ingestion and API payloads.
  • Markdown Table (`.md`): GFM-compliant markdown syntax for developer documentation and notes.
  • Clipboard TSV: Tab-separated values copied directly to the clipboard for 1-click pasting into Excel, Google Sheets, or Notion.

11. Deployment Runbooks & Self-Hosting

DevOps & Cloud

The full-stack application supports both self-hosted Docker environments and managed Vercel serverless deployments:

Bash — Self-Hosted Docker Deployment
# 1. Clone repository (downloads clean Next.js app without demo files) git clone https://github.com/ederaefe/Visiolog.git cd Visiolog # 2. Configure environment variables cp .env.example .env.local # 3. Launch full stack with Docker Compose docker compose up -d