All projects
Automation & ReconciliationIndustry Project · NoBrokerHood

ScrapeForge

ERP-migration automation: reliable multi-year data extraction from legacy systems, then reconciliation you can trust.

My roleDesigned and implemented the system end-to-end, including the AI logic, backend, frontend, data processing and automation workflows.

Problem

Problem

When housing societies move to NoBrokerHood, years of financial records sit in legacy ERPs. Extracting them by hand is slow, and without systematic reconciliation nobody can be sure the migrated books agree.

System pipeline

  1. Source ERP
  2. Automation
  3. Extraction
  4. Normalise
  5. Reconcile
  6. Readiness

Approach

Validation is part of the system, not an afterthought.

Automates extraction of multi-year financial reports from legacy ERPs with checkpoints and count validation, then reconciles them against the new system. The workflow keeps inputs, decisions and exceptions visible so that a reviewer can understand how the output was reached.

My contribution

  • Browser-session and report extraction workflow
  • Financial-year, date, pagination and checkpoint handling
  • Canonical data model, matching and difference classification
  • Readiness scoring, progress events and recovery tools

Try the demo

ScrapeForge: interactive demo

Simulated legacy ERP → real filtering, pagination, extraction, normalisation, de-duplication and reconciliation.

  • Computed
  • Simulated visual
  • Local model
Session ready

Create a synthetic browser session.

ledgerlegacy.demo / receipts
Reports Receipts Reconciliation
Receipts report2026-01-01 → 2026-09-30

Ready for a synthetic extraction run.

◇

Choose a date range to run real filtering, de-duplication and reconciliation.

This interactive demonstration is a simplified, synthetic representation inspired by an enterprise project. Production systems, company data and proprietary implementation details are not publicly exposed.

Result

A result that can be inspected.

The portfolio demo filters and paginates synthetic records, removes exact duplicates, reconciles two datasets and derives its readiness grade from the computed match rate.

Technical architecture

Stage-by-stage system design

  1. 01
    Source ERP · One browser, one session

    Legacy ERP session → Adopted browser session

  2. 02
    Automation · Browser automation

    Report + date range → Loaded, filtered report

  3. 03
    Extraction · Count-validated extraction

    Paginated table → Raw dataset + checkpoints

  4. 04
    Normalise · Canonical financial model

    Raw datasets → Canonical records

  5. 05
    Reconcile · Explainable reconciliation

    Canonical records (old vs new) → Matches + difference categories

  6. 06
    Readiness · Migration readiness scoring

    Reconciliation results → Readiness report

Tech stack

Tools used in the production system

  • Playwright · Chromium
  • FastAPI
  • Pandas · RapidFuzz
  • WebSockets · noVNC

The portfolio demo itself is static TypeScript running locally in the browser with synthetic data and no API key.