The durable home for how Sorta is being built, what is finished, and what comes next.

· Combined website release

Green redesign and AI-readable public content

The owner approved publishing the green redesign and public-readability improvements together. The earlier local-only notes below describe preparation milestones; this entry records the combined release checks and supersedes their open preview issues.

The production preview now runs the actual hosting build using Cloudflare’s local runtime. Its local settings are resolved from the project folder, without copying credentials into the build or enabling remote resource bindings. All existing declared application-package versions are retained; a pinned development-only preview tool was added.

The earlier security warning was reviewed and addressed with the framework’s same-origin protection for server-function requests. Normal website requests pass; cross-site and unverified action requests are rejected. Public HTML, the sitemap, the content index, and the separately authenticated training-agent API are outside this origin check. Existing sign-in, database permissions, and private storage protections remain in place.

Release verification: all 232 tests pass, including 15 public-content checks against the production build without JavaScript and 11 origin-protection tests. Read-only production-preview checks also confirm same-origin public-data requests succeed and anonymous requests for private Sorta records are rejected without exposing image URLs. The build, TypeScript check, and targeted code checks pass. No image records, annotations, review decisions, datasets, or training runs were changed by this release.

Published to ftggcig.com on 6 September 2026 through the existing GitHub and Lovable connection, without Lovable chat prompts. All 15 live public-content checks pass: ten readable public pages, the sitemap and plain-text content index, protected-page indexing rules, and a real missing-page response. The live stylesheet contains the approved green accent, and MyMoney Journey includes actual milestone content in its initial HTML. Live read-only checks also confirm the origin protection and rejection of anonymous requests for private images. No browser visual or authenticated editor testing was performed during this release.

One separate web-reading service still rejected the domain with a retrieval safety error while direct public HTTP checks passed. The cause remains unconfirmed; this is not evidence that every AI provider can retrieve the site or that a site firewall caused the failure. No security protection was disabled. Search indexing and citations remain controlled by each service, not guaranteed by these changes.

Current project status

The foundation and annotation workflow are built. Phase 3 focuses on reviewed data, versioned datasets, and a local training workflow. A trained recognition model is not yet available on this website.

The dated records below preserve decisions and corrections from each stage of the build. Earlier plans and limits are historical, not a statement of today’s capabilities. This website-readability update does not import images, change review decisions, or run training.

· Public-site implementation

Public content delivered directly to readers

Following the audit below, the owner authorized targeted improvements for AI readers. The existing website framework and green design are retained. MyMoney Journey now waits for its public updates before rendering, so titles, dates, and entry text are available in the initial HTML without JavaScript. Editing still requires sign-in; an initial data-loading failure is not presented as an empty collection of updates.

One explicit list of ten public pages now drives the sitemap, canonical page addresses, and structured descriptions of the organization, website, projects, and breadcrumbs. The supplementary AI-readable content index links to the same public pages and states Sorta’s actual development stage. The sitemap includes project overviews, Documentation, Journey, Privacy, and Support. It does not invent content-update dates. Historical notes are distinguished from current status.

Public pages explicitly permit indexing; login, private, and unmatched routes do not inherit that permission. Private image URLs, annotations, review records, and training assets are excluded from the new indexes. Existing authentication, database access rules, and crawler rules are unchanged. Search discovery and model-training permission remain separate choices; metadata is not a substitute for access control.

Verification in the development preview: 221 tests pass, including 15 opt-in HTTP checks without JavaScript and 25 new metadata and Journey rendering tests. All ten public pages return readable content, one canonical address, indexable metadata, and valid JSON-LD. MyMoney’s actual milestone entries are present. Login and workspace responses remain non-indexable; an unknown page returns HTTP 404. The production build, TypeScript check, and targeted lint checks pass. No browser-interaction or authenticated review testing was performed.

An additional production-preview check was blocked by the existing local preview setup: it requested a server file under dist/server, while the hosting build emits a Cloudflare worker under .output/server. That preview returned errors before the application ran. The production artifact still needs a check in its supported hosting runtime; development-preview results are not a claim that this deployment check passed.

Provider-specific crawl access and indexing still require checks after publication. The llms.txt format is supplementary guidance, not a guarantee that every service will read or cite the site. No firewall protection was disabled and no search or model-training submission was made.

Prepared locally; this entry does not indicate that the changes have been published to ftggcig.com.

· Public-site audit

Making public project knowledge readable by AI

The owner requested that AI agents be able to read and reference the public FTGGCIG website. Direct checks of the live site, without running JavaScript, confirmed readable content on the homepage, Why, What, Who, Sorta overview, Sorta Documentation, and MyMoney overview. The site is not universally an empty JavaScript shell.

At audit time, MyMoney Journey was an exception: its initial HTML contained “Loading updates…” rather than the entries. The sitemap listed only the four main pages and omitted the public project pages. Canonical URL tags and structured data are also missing from the checked pages. These findings led to the implementation entry above; they describe the deployment before those improvements.

The live crawler rules allow access. One web-reading service nevertheless failed to retrieve the site while ordinary public requests succeeded; the reason remains unconfirmed. Hosting crawler-access logs should be checked before attributing this to JavaScript or changing security settings.

Search visibility and model-training permission are separate choices, as described in the official OpenAI crawler documentation. Readability improves the opportunity for discovery and citation, but does not guarantee either. Private images, reviews, and workspace data remain protected; no crawler policy, application behavior, or access control was changed.

This audit record is local and awaits publication with the approved website changes.

· Design update

A bolder look, the same careful workflow

Before resuming Phase 3, the site’s visual direction was updated using bored.com as a reference: bold type, crisp outlined cards, compact corners, offset shadows, and a subtle dotted background. Bright green replaces the reference’s yellow accents at the project owner’s request. Our name, content, and identity remain our own.

The shared header and footer, home page, Why / What / Who pages, project overviews, project tabs, workspace navigation, and this Documentation tab carry the new style. Sorta AI Workspace and MyMoney download links retain their existing destinations. The Workspace tab stays highlighted while using its child pages.

Annotation pages keep a neutral, undotted background. Category colors, masks, image rendering, zoom and drag controls, review decisions, private storage, and training eligibility are unchanged. No images were imported, no records modified, and no dataset or model version created by this design work.

Navigation wraps on smaller screens, mobile controls retain accessible names and expanded state, and keyboard users have a skip-to-content link and visible focus outlines. System fonts avoid adding a third-party font request.

Verification: all 181 automated tests pass, including four new navigation checks; the production build and TypeScript check pass. Targeted code checks also pass (legacy formatting differences were excluded). Public-page response checks pass. Existing framework deprecation warnings remain. Browser interaction testing and authenticated manual review were not performed as part of this design change.

A pre-existing framework warning about server-function CSRF protection also surfaced locally. It requires a separate security review before publication; this styling change neither alters security behavior nor claims to resolve that warning.

Prepared locally for visual approval; this entry does not indicate publication to ftggcig.com.

Phase 1 — Foundation, taxonomy, and data model

Historical foundation record, followed by later phase updates. Phase 1 was a deliberately narrow vertical slice: authentication, private image intake with persisted metadata, the versioned V1 taxonomy, and independent review-state dimensions.

Principles and guardrails

  • Build in approved phases. No annotation canvas, dataset import, training, inference, or model dashboards until each is separately authorized.
  • Classification decisions and segmentation decisions are tracked separately and are never collapsed into a single status.
  • Residual means a visually identifiable item belonging to none of the other seven categories. It never stands in for uncertainty, blur, or difficulty.
  • Hazardous is assigned only from visible identity or an approved rule — unknown contents are never sufficient.
  • Tags are optional descriptive metadata. They are never semantic classes and never part of a model target.
  • Source images stay private and are never published on this site.

Foundation completed

  • Authenticated workspace with server-side and database-level access control.
  • Image intake: select an image, store it privately, and persist its metadata (filename, type, size, storage location, provenance, uploader, timestamps, review status).
  • Recent uploads listing with short-lived signed previews, scoped to the uploader.
  • Versioned taxonomy V1 with exactly eight categories: Organics, Fiber, Plastic, Metal, Glass, Inert/Mineral, Hazardous, Residual.
  • Audit rows written on image registration and on review-status change.

Architecture, data, and storage

  • TanStack Start with file-based routing, React, TypeScript, and the site's existing Tailwind/shadcn design system.
  • Canonical routes: /projects/sorta (public overview), /projects/sorta/workspace (authenticated), and /projects/sorta/documentation (this page). Legacy /waste-ai, /projects/sorta/model, and /projects/sorta/metrics redirect to the new locations.
  • A normalized, prefixed schema separates taxonomy versions, categories, sources, images, instances, annotations, tags, review events, and settings, with provenance and soft-deletion preserved.
  • A dedicated private storage bucket holds source images under per-user, collision-safe paths. Storage access is written through an abstraction layer so S3-compatible storage can replace it later.

Verification status

  • Automated tests cover taxonomy stability, independent state dimensions, upload metadata validation, tag normalization, user-scoped paths, and upload cleanup.
  • Typecheck and production build pass.
  • Access rules, database constraints, and storage policies are exercised through the running application rather than by automated integration tests.

Taxonomy V1.0 — approved

All eight Version 1 categories (Organics, Fiber, Plastic, Metal, Glass, Inert/Mineral, Hazardous, Residual) now have approved definitions with structured include, exclude, caution, and note rules stored alongside them. No further semantic classes exist.

Classification precedence is fixed:

  1. Hazardous / special handling, only from visible identity or an approved rule — never from uncertainty.
  2. Otherwise the defensible primary material category.
  3. Otherwise Residual, and only when positively identified.
  4. Otherwise a review state — needs review, ambiguous, unidentifiable, or excluded. Review states are never semantic categories or training labels.

A credible but visually unconfirmed hazard routes to needs review with an optional potential-hazard flag. The acceptance threshold for routing low-confidence predictions to review is configuration, initially 0.80, and is not an inference implementation. The full approved specification is kept in the repository at docs/waste-taxonomy-v1.0.md.

Phase 2 — TACO 10-image pilot (complete)

Phase 2 is authorized only as a tightly bounded pilot: exactly ten images from the official pedropro/TACO dataset, pinned to a recorded source revision and checksum in a committed manifest so the import is reproducible and idempotent.

  • Original COCO image ids, filenames, dimensions, licenses, categories, and polygon segmentations are preserved as source provenance. Imported annotations are never destroyed or silently overwritten — corrections create a new current human revision and keep the imported record.
  • TACO categories are mapped to Taxonomy V1.0 through versioned mappings with a status of direct, conditional, manual review, or excluded. A mapping only ever suggests a class; it never approves an instance.
  • Everything imported stays unapproved. Category approval and mask acceptance are separate human actions, and an image only becomes reviewed when the server confirms every object has been addressed.
  • Ambiguous, unidentifiable, and excluded remain review states, never semantic categories, and never become training-eligible.
  • A reviewed-annotation COCO export is available for reviewed images and training-eligible instances only. It is not Dataset v001.
  • The annotation editor, editor queue, image dataset, taxonomy mapping screen, and import all live inside the private, authenticated workspace. TACO images are never shown publicly and are served only through short-lived signed URLs.

Status (superseded — see the 4 September 2026 Phase 2 closure entry below): at the time of this entry the ten pilot images were imported and the review tooling was live, with the human review of all ten still outstanding. That review, and final human acceptance QA, are now complete and Phase 2 is closed; the full TACO import remains deliberately deferred and separately authorized.

Update — 3 September 2026: a source-agnostic annotation platform

TACO is the initial bounded pilot source, not the only source of data. The annotation platform is source-agnostic: authenticated manual uploads are first-class work items that enter the same editor queue and the same annotation editor, and their reviewed annotations count toward the reviewed COCO export.

  • Upload-to-review workflow: upload an image in the private workspace, then open it directly in the annotation editor, draw missed objects, choose one of the eight Taxonomy V1.0 categories, approve the category and accept or correct the mask as separate actions, add notes and optional tags, and complete the image only when the same server-side guard confirms every object is addressed.
  • Provenance rules: uploaded images and human-created instances keep their own manual-upload provenance and are labelled as such in the dataset, the editor, and the export. They are never presented or exported as TACO data.
  • Canvas interaction: zoom is button-only, through compact controls anchored to the bottom-right of the image viewport. Wheel, trackpad, and pinch gestures never zoom the canvas. Panning requires a deliberate press-and-drag; a click without movement never shifts the image, and panning never starts while drawing or moving a vertex.
  • TACO remains hard-limited to exactly ten images in the manifest and again in the server-side importer. The full TACO import is still blocked pending completion of the pilot review and the resulting annotation UX review.

Phase 2 remains in progress.

Update — 3 September 2026: workspace information architecture

The private workspace is organised into four tabs: Dashboard, Editor, Image dataset, and Taxonomy mapping. The former review queue is now the Editor tab; its old address redirects there.

  • Dashboard is a high-level operational view over every accessible active image: total images, queued (not started), in review, reviewed, excluded, completion percentage, total objects, training-eligible objects, and a source breakdown of TACO versus manual uploads. Counts are computed from distinct image and instance rows, so joins cannot inflate them. The TACO pilot appears only as a compact source metric and import action (superseded by the 3 September 2026 entry below).
  • Editor holds queued annotation work. It shows imported and in-review images from every source by default, with status and source filters, and each card carries a thumbnail, source, status, object counts, upload time, and an Open editor action.
  • Image dataset is the complete library, regardless of review status. It carries the manual upload form, filename search, source and review-status filters, and sorting by most recently uploaded (default), oldest uploaded, most recently reviewed, or filename. Every row opens in the editor.
  • Workflow: upload in Image dataset, open the new image directly in the editor, annotate it, and complete it through the unchanged server-side review guard.
  • Timestamps: created_at is the authoritative upload or import time; reviewed_at is the latest successful review completion.
  • Reviewed images can be deliberately reopened from Image dataset for correction. Any substantive annotation, category, or mask change to a reviewed image returns it to in review, so stale reviewed data cannot stay training-eligible; the historical reviewed timestamp and audit trail are preserved until the image is completed again.
  • TACO remains one source among others and remains capped at exactly ten images.

Verification snapshot (3 September 2026): the implementation passed 79 automated tests, typecheck, the production build, and lint on the changed files. Live post-migration verification showed 12 active images in total — 10 TACO pilot images and 2 manual uploads — with 5 reviewed and 7 queued, and reviewed_at populated for every reviewed image. These counts are a dated snapshot only; review progress continues to change.

Phase 2 remains in progress.

Dashboard cleanup — 3 September 2026

A presentation-only cleanup made the workspace Dashboard source-neutral. No routes, data, review states, images, annotations, or authentication behaviour changed.

  • The duplicate "Review queued images" and "Open image dataset" buttons were removed from the Dashboard, because the persistent Workspace tabs directly above them are the canonical navigation.
  • The standalone "TACO source — bounded 10-image pilot" box, with its import and export buttons, was removed from the Dashboard: TACO is one image source, not the identity of the workspace.
  • TACO stays represented neutrally as one ordinary line in the Dashboard source breakdown alongside manual uploads, and remains hard-capped at exactly ten images.
  • Reviewed COCO export moved to the Image dataset tab as a general, cross-source dataset action, keeping the same authenticated export behaviour with loading, error, and success feedback.
  • The idempotent TACO importer and its cap remain in the codebase for maintenance and reproducibility; because the pilot import is already complete, that action is no longer permanently displayed in the interface.
  • Rationale: reduce redundant navigation and keep the Dashboard source-agnostic.

Verification snapshot (3 September 2026): automated tests, typecheck, the production build, and lint on the changed files all passed. Live verification confirmed the data was untouched — 12 active images (10 TACO pilot images, 2 manual uploads), 5 reviewed and 7 queued, with 32 annotation instances. These counts are a dated snapshot only.

Phase 2 remains in progress; no later phase was started.

Missed-object drawing feedback — 3 September 2026

The annotation editor's "Add missed object" drawing tool now gives immediate visual confirmation of every click. No images, annotations, review states, routes, or authentication behaviour changed.

  • Each successful click on the canvas renders a high-contrast vertex marker at the registered point, readable over both light and dark images.
  • Registered points are connected by visible draft edges: an open dashed path until three points exist, and only then a lightly filled closed shape.
  • The first point is drawn distinctly so the polygon's starting vertex is obvious, and a live point count sits beside the draw controls.
  • Finish polygon stays hidden until at least three points exist; Cancel and Escape behave as before, and drawing feedback is non-interactive so one click can never register two points.
  • Draft markers, edges, and the point count reset when the polygon is finished or cancelled, when Escape is pressed, and when another image is opened.
  • Draw mode continues to take precedence over panning; button-only zoom and click-drag panning outside draw mode are unchanged.
  • Rationale: users must be able to confirm that each click was registered while tracing an object, instead of inferring it from a shape that only appears later.

Verification snapshot (3 September 2026): 84 automated tests, typecheck, the production build, and lint on the changed files all passed, including new regression tests covering ordered point appending, the three-point finish rule, and draft reset. These counts are a dated snapshot only.

Phase 2 remains in progress; no later phase was started.

Automatic contain-fit in the annotation editor — 4 September 2026

Every image now opens fully visible and centred in the editor viewport. No images, annotations, review states, routes, or authentication behaviour changed.

  • Opening the editor, and navigating to another image, automatically fits and centres the whole image using true "contain" geometry from the measured viewport size and the image's original pixel dimensions.
  • Portrait, landscape, and square images are all shown in full with no cropping and no stretching; leftover space is centred as letterboxing.
  • The Fit control (and f) restores exactly this centred contain baseline rather than a top-left zoom of 1.
  • The fitted view is 100% in the interface, and the zoom buttons scale relative to that baseline within the existing clamped range.
  • The fit is recomputed when the viewport resizes; zoomed, panned, drawing, and vertex-editing work is never reset underneath the user, and stored polygon coordinates are unchanged because only the display transform moves.
  • Manual uploads whose pixel dimensions are backfilled after load fit as soon as those dimensions become available.
  • Fully compatible with button-only zoom, deliberate drag-to-pan, and the per-click polygon vertex feedback; wheel, trackpad, and pinch zoom remain disabled.

Verification snapshot (4 September 2026): 95 automated tests, typecheck, the production build, and lint on the changed files all passed, including new contain-fit regression tests for orientation, invalid input, centred offsets, fit-as-100% semantics, and coordinate mapping after fit, zoom, and pan. A dated snapshot only.

Phase 2 remains in progress; no later phase was started.

Update — 4 September 2026: Phase 2 status audit

Live status audit of the bounded ten-image pilot, dated 4 September 2026. This entry supersedes earlier wording that described the ten-image human review as unfinished.

  • 12 active images: 10 TACO pilot images and 2 manual uploads.
  • All 12 images are REVIEWED; the queued count is 0.
  • All 10 TACO pilot images are REVIEWED.
  • 34 active instances: 32 imported and 2 human-added.
  • 33 training-eligible instances.
  • 1 active false-positive instance is retained in audit and provenance and is not training-eligible.
  • Every active instance is HUMAN_VERIFIED for classification; 33 masks are MASK_ACCEPTED and the false positive is MASK_CORRECTED.
  • The review audit covers category selections and approvals, mask acceptance, 2 polygon revisions, 2 missed-object additions, false-positive marking, and image completion.
  • 60 current TACO mappings remain versioned: 34 DIRECT, 20 CONDITIONAL, and 6 MANUAL_REVIEW.

Conclusion (superseded by the 4 September 2026 Phase 2 closure entry below): the bounded ten-image Phase 2 implementation and the pilot review were complete in the database, with final human acceptance QA outstanding at the time of the audit.

The full TACO import is deliberately deferred. It is not part of closing this user-authorized ten-image pilot and requires separate authorization after pilot acceptance. Dataset v001, training, model work, and inference remain Phase 3 and later, and remain unauthorized.

Automated verification is already green: 95 tests, typecheck, the production build, and changed-file lint. An independent live visual browser pass of the latest contain-fit behaviour could not be completed in this session because the preview required authentication; that check is not claimed as passed.

Update — 4 September 2026: Phase 2 complete

Phase 2 is formally COMPLETE. This entry supersedes all earlier wording that described Phase 2 as in progress or listed outstanding Phase 2 work. Documentation only: no application behaviour, routes, authentication, database records, images, annotations, or review states changed with this entry.

  • The final human acceptance QA was completed by the project owner after the latest editor UX changes (contain-fit, button-only zoom, drag-to-pan, per-click vertex feedback, mask correction, and reviewed-image reopening).
  • The manual spot-check of the reviewed COCO export was completed — files, polygons, class labels, image dimensions, and provenance.
  • Any Phase 2 defects requiring closure have been addressed. There is no remaining Phase 2 work.
  • The audited live snapshot is preserved: 12 active images, all 12 reviewed; 10 of 10 TACO pilot images reviewed; 34 active instances; 33 training-eligible; one retained false positive; 60 current versioned TACO mappings.
  • Automated verification remains green: 95 tests, typecheck, the production build, and changed-file lint.
  • The previously documented authentication limitation on the independent browser pass remains historical context only; the owner's completed human acceptance supersedes it as the closure decision.
  • The full TACO import remains separately deferred and requires explicit authorization. It is not an unfinished item of this bounded ten-image pilot.
  • Dataset v001, training, model work, and inference are Phase 3 and later, and have not begun.

Phase 3 foundation — 5 September 2026

The Phase 3 foundation is built: reproducible dataset snapshots, a local training agent, and a model registry. It is deliberately unused. Dataset v001 has not been created, no training run has been started, and Model v001 does not exist. No additional TACO images were imported; the ten-image cap holds. Inference remains Phase 4 and has not begun. Nothing was deployed.

  • Eight new tables record dataset versions and their frozen image and instance membership, training runs, evaluation runs, model versions, metrics, and an append-only audit trail. All are protected by row-level security for authenticated team members only, with no anonymous access and no deletion.
  • Lifecycle guards live in the database as well as the application: a snapshot moves Draft → Frozen → Archived and a frozen snapshot is immutable; a training run ends Succeeded, Failed or Cancelled and never reopens; a model reaches Champion only through an explicit human promotion, and at most one champion can exist at a time.
  • Two private storage buckets hold snapshot manifests with their COCO export, and model artifacts. The training agent receives only short-lived signed URLs; no service key ever leaves the server.
  • Splits are deterministic and leakage-safe. The unit is the group, not the image: images sharing an exact SHA-256 checksum, or a manual grouping, always land in the same split. Ordering derives only from the seed and the group key. Near-duplicate (perceptual) clustering is not implemented, and that limitation is stated rather than assumed away.
  • Every snapshot freezes a canonical manifest: image and instance identifiers, the current annotation identifiers and their geometry checksums, taxonomy and annotation versions, source provenance, split assignments, per-class counts, the creator, the date, the notes, the confirmed policy, and the export checksum. The manifest checksum is reproducible on both sides.
  • The readiness gate is configurable and separates blockers from warnings. Creating a snapshot requires typing a confirmation, echoing back the exact readiness hash that was displayed, and giving a written reason for any accepted blocker. Every proposed threshold — the 70 / 15 / 15 ratios, the fixed seed, the class and checksum requirements — is a proposal awaiting operator confirmation, not an approved rule.
  • The local training agent lives in the repository under training-agent/ and runs separately from the website. It signs in with the operator's own workspace login, rejects a service key, verifies every downloaded file against the manifest before use, never re-splits, prefers MPS then CUDA then CPU while recording each fallback, and keeps the Mask R-CNN architecture behind a swappable adapter.
  • Two new private workspace tabs display the work: Dataset versions (readiness audit, split proposal, dry-run preview, gated create and freeze) and Training & models (agent setup, frozen datasets, run history, model registry, baseline scorecard).

Live readiness on this date: 12 active images, all reviewed, 33 training-eligible instances. Class coverage is FIBER 2 images / 3 instances, GLASS 4 / 4, HAZARDOUS 1 / 1, METAL 4 / 7, ORGANICS 1 / 1, PLASTIC 7 / 17; INERT_MINERAL and RESIDUAL have no eligible instance at all, and two images still have no stored checksum. The gate therefore reports blockers, which is the correct outcome: a snapshot taken today could not honestly claim an eight-class baseline.

Verification snapshot (5 September 2026): 124 automated tests, including 29 new Phase 3 tests for checksums, canonical serialization, deterministic and seed-sensitive splits, duplicate confinement, readiness blockers and warnings, lifecycle transitions and manifest stability, plus typecheck, the production build, and changed-file lint. The Python agent tests are CPU-only and could not be executed in this environment because pytest is not installed there; that check is not claimed as passed. The full written record is in docs/waste-ai-phase-3-dataset-and-training.md.

Phase 3 correction and hardening — 5 September 2026

The “Phase 3 foundation complete” status recorded earlier the same day is retracted. An audit found defects serious enough that the label was not earned. They are listed below with what was done about each. The corrected foundation is now believed sound, but it is still an unused foundation: no dataset version, training run, evaluation run or model version exists, no images were imported, no review state changed, nothing was deployed, and Phase 4 inference remains out of scope.

  • The readiness audit reported an empty corpus. The candidate query never resolved an instance's category, so every instance failed the eligibility test and the page implied there was nothing to train on. Categories are now resolved before the test, and the audit reports the real corpus: 12 reviewed images and 33 eligible instances. Instances that are deleted, marked false positive, or have no current geometry are excluded, as intended.
  • Snapshots were not genuinely immutable. A frozen snapshot recorded which annotations belonged to it but re-read their geometry from the live tables at export time, so later editing would silently change a “frozen” dataset. Each membership row now stores the exact geometry itself alongside its checksum; freezing and the COCO export read only that stored copy, and every checksum is recomputed and compared before anything is written. A single mismatch aborts the freeze.
  • The manifest was partly placeholder. It could be written with a blank readiness hash, empty counts, a fresh timestamp instead of the recorded annotation watermark, and the snapshot code standing in for its name. It is now built entirely from the stored snapshot record and its frozen membership, and validated against its schema before upload; missing provenance is an error, not a blank field.
  • Snapshot creation could leave half-written state. Creation now runs as one database transaction through a single stored procedure, so a failure leaves no partial snapshot behind. Freezing validates everything before uploading and removes any object it just wrote if the final step fails.
  • Stored files were not scoped to their owner. Manifests, exports and model artifacts now live under a per-user prefix, and the storage rules check that prefix, so one signed-in operator can no longer read or write another's files. Every Phase 3 table carries ownership-based access rules, and every agent endpoint filters and validates by the signed-in user: uploads only into the caller's own successful run, registration only of an artifact that actually exists at the expected size, and new models only ever as candidates.
  • Membership rules misbehaved on deletion. The guard referred to a row that does not exist during a delete, which would error. Draft membership can now be corrected or rolled back by its owner; frozen and archived membership stays immutable.
  • The training agent could have mislabelled objects. It inferred the class list from whichever classes happened to appear, and matched images by guesswork. It now reads all eight classes from the manifest in their defined order, resolves each image by its recorded identifier, and refuses to label a real object as background. Checksum verification can no longer be skipped, an access token may be supplied directly instead of a password, the saved session file is owner-only, and the run now produces and uploads its log, checkpoints, configuration, metrics and evaluation report, with evaluation results recorded back in the workspace.

Verification on this date: 138 automated website tests (14 new ones covering eligibility, frozen geometry integrity, export determinism, manifest population and owner-scoped paths), typecheck, production build and changed-file lint all pass. The agent's 23 Python tests pass, including a new contract test whose fixtures are generated by the same website code that writes a real snapshot; it proves a non-empty training split, correct image matching and a correct non-background label for all eight classes. Both protected workspace pages were opened and checked. A live database check confirms 12 images, all reviewed, 34 instances and 33 eligible, unchanged, and every Phase 3 table still empty. The full written record is in docs/waste-ai-phase-3-dataset-and-training.md.

Phase 3 pre-import cleanup — 5 September 2026

A small, focused cleanup made before any data import. Nothing was created, imported or deployed: there is still no dataset version, training run, evaluation run or model version, and Phase 4 inference remains out of scope.

  • The training tool had two copies of the same two commands. Duplicate definitions for creating and updating an evaluation shadowed each other; one correct pair remains.
  • Evaluation scores would not have been recorded. The tool read the scores from the wrong place in its own report, so a real evaluation would have saved no numbers at all. It now reads the actual report and saves each score (overall accuracy and its variants) for both box and mask results. A score that could not be computed is recorded with the reason instead of a misleading zero.
  • Intermediate checkpoints were never uploaded. Registering a model now uploads every per-epoch checkpoint alongside the final model, its configuration, metrics, log and evaluation report, with filenames checked and duplicates impossible.
  • Cached Python build files were committed. They are removed and now ignored.
  • The all-or-nothing snapshot step is now pinned by a test. Its presence and exact name were confirmed directly in the database.
  • Measurement permissions were too loose. A signed-in operator can now only attach a measurement to a snapshot, run or model they own, matched to the kind of measurement.

Verification on this date: 141 automated website tests, 28 agent tests (five new ones covering both evaluation report shapes and checkpoint discovery), typecheck, production build, changed-file lint and a Python compile check. A live database check confirms 12 images, all reviewed, 34 active instances and 12 reviews, unchanged, with every Phase 3 table still empty. The full written record is in docs/waste-ai-phase-3-dataset-and-training.md.

Image library expanded to 200 pictures — 6 September 2026

The authorised expansion added 188 more unique TACO pictures, bringing the library to 200. Pictures were chosen deliberately: first those showing the waste types we had none of, then the rare ones, then for the widest spread of types, with near-identical pictures rejected automatically. Eleven look-alikes were skipped.

A live check confirms 200 pictures, every one distinct, 1,355 marked-up objects (34 already checked by a person, 1,321 waiting), and exactly 200 stored files with no gaps, no leftovers and no mismatched sizes or shapes.

  • Fixed: some pictures showed “0 objects”. The workspace could only read the first thousand objects at a time, so anything past that looked empty. It now reads them all, in pages.
  • Fixed: some previews did not appear. Picture links are now requested in batches rather than one at a time, so every thumbnail loads.

Checked in the running workspace: correct object counts, every preview visible, no errors, and an imported picture opening in the editor with its outlines intact.

Checking a picture in one action — 6 September 2026

With 188 pictures waiting, approving each object separately was the slowest part of the work. A picture can now be confirmed with a single action, while the judgement stays with the person: the objects are numbered on the picture and in the list, each shows where its suggested waste type came from, and any type that only applies under a condition shows that rule in plain words.

Before confirming, the panel spells out exactly what will be recorded — how many suggested types will be accepted, how many of those are conditional, how many choices you made will be approved, how many outlines will be accepted, and how many objects are being kept out of training. Anything unresolved is listed as a numbered item you can click to jump straight to that object.

  • Nothing is recorded unless the whole picture passes: a broken outline, or an object whose source type needs an individual decision, stops the action and leaves the picture exactly as it was.
  • Your own choices win. A type you picked yourself, an outline you corrected, an object marked as not really there, and anything set aside as unclear all keep their decision and stay out of training.
  • If the picture changed since you opened it, the action is refused and asks you to reload, so two people cannot overwrite each other.
  • Confirming twice records nothing the second time, and every completed picture leaves a single audit entry with its counts.
  • Outlines are checked strictly: an outline must match the picture's real size, stay inside it, have at least three distinct corners and enclose a real area.
  • A waste type that is no longer part of the current eight-type list, or a source rule that has been changed since the pictures were brought in, stops the action and asks for an individual decision.
  • Everything is recorded together or not at all: the picture is re-checked at the last moment, and if anything moved while the action was running, the whole thing is abandoned and nothing is written.
  • If you are part-way through drawing a missed object, or editing an outline you have not saved, the action is unavailable and says so — it becomes available again once you save or clear that drawing.

Verified on 6 September 2026 against throwaway test pictures put through the real, signed-in path and then removed: signed-out attempts refused, an out-of-date attempt refused, objects needing a human decision blocked with a reason each, a clean picture confirmed (one suggested type accepted, two of your own choices approved, three outlines accepted) and a repeat click recording nothing. The library was confirmed identical afterwards — 200 pictures, 12 checked, 188 waiting, 1,355 objects with 1,321 still unreviewed, 34 approved.

For the record, the waste types shown on the 1,321 waiting objects are suggestions carried in with the pictures, not confirmed types: 573 residual, 369 plastic, 108 glass, 101 fibre, 54 metal, 5 inert mineral and 111 with no suggestion at all. By confidence, 400 are direct suggestions, 810 apply only under a condition a person must look at, and 111 need an individual decision. None of this counts as a checked result.

Fixed: conditional waste types blocked every one-action check — 6 September 2026

Reported on the day: confirming a picture whose types and outlines all looked correct failed with “Nothing was changed. Source mapping needs an individual category decision.” repeated twice, with no indication of which objects were at fault.

Cause. Every source rule that applies only under a condition is also marked “a person must look at this”. The editor allowed the one-action confirmation, but the save routine treated that mark as “must be decided object by object” and refused the whole picture. That disagreement affected all 20 conditional rules, covering 810 of the waiting objects. The number of affected pictures was not established.

Fix. The mark now means what it says: a person must attest, and the deliberate one-action confirmation is that attestation. A conditional type may therefore be accepted by that click, provided it is the current rule, unchanged since the pictures were brought in, points at a live waste type, and has its rule shown on the object. Everything else stays blocked exactly as before: rules needing an individual decision, rules with no suggested type, rules excluded, and — kept deliberately strict — a plain direct rule that unexpectedly carries the “must look” mark. Outline checks, current-type checks, the out-of-date reload guard, your own choices, corrected outlines and set-aside objects are all untouched. The messages now name the object number and the real reason instead of repeating one generic line.

Evidence. The full test suite passes (177 checks, 13 files), including 26 for this screen with new cases built from the real rule data: conditional-with-mark accepted, direct-without-mark accepted, direct-with-mark blocked, must-look blocked, no-suggested-type blocked, changed-since-import blocked, and per-object messages. The live check itself was exercised against the real database with throwaway records inside a transaction that was rolled back: two conditional-with-mark objects completed in one action (two types accepted, two outlines accepted, two training-eligible); a mixed picture with one unresolved object was refused with a single message naming that object and wrote nothing at all; and a direct rule carrying the mark stayed refused. The library was confirmed identical afterwards — 200 pictures, 12 checked, 1,355 objects, 34 approved, 1,357 outlines, 60 current rules — with no test records left behind. Nothing was published, no pictures were brought in, no training was run and no dataset version was created.

What Model v001 has to be

Model v001 is a carefully measured, repeatable first model aimed at the best practical recognition and outlining performance across our eight waste types with this library. It is not a demonstration. Training only begins once all of the following hold:

  1. A person has checked every outline that goes into the training set.
  2. Each waste type has enough examples, with the rare ones deliberately accounted for.
  3. Similar pictures are kept together on one side of the split, so nothing the model trained on reappears in the test.
  4. The approved way of splitting the data is recorded and locked with the set.
  5. Results are measured on pictures held back from training, for both boxes and outlines, with a breakdown of mistakes per waste type.
  6. Those findings feed back into the mark-up and category work before any model is promoted.

No training set, training run, evaluation or model has been created yet, and putting a model to work inside the site remains a later phase.

Direct local development setup — 6 September 2026

The private code repository is now named mlinton-ftggcig/for-the-greater-good-community. Lovable detected the rename automatically and confirmed that its main branch remained connected and in sync. The repository owner, privacy settings and complete change history were preserved.

GitHub Desktop was installed and connected with the operator's approval. A local copy containing all 578 commits was verified against the latest review fix, 98d529a. Development can now use direct file edits without sending Lovable chat prompts; changes reach GitHub and Lovable only after they are committed and pushed.

This setup also corrects the explanation of the conditional-mapping bug above: the editor allowed confirmation while the save routine rejected it. No image data, reviews, model training or deployment settings were changed. These documentation edits were prepared locally; the setup itself did not publish the website or rerun the full application test suite.

Documentation practice

Every material change to Sorta receives a dated entry here covering its scope, the rationale and resulting behaviour, the affected workflow, routes and data, how it was verified, and what remains outstanding along with the current phase boundaries. This Documentation tab is the living project record.

Historical Phase 2 handoff — superseded by the Phase 3 updates above

The following limits and sequence were recorded at the Phase 2 handoff. They are retained for context, not as the current import cap or current implementation status.

  • Not built yet: dataset snapshots, the full TACO import, training, inference, active learning, and metrics dashboards.
  • Access is "any authenticated team member"; roles come later.

Phase 1 and Phase 2 are complete. Phase 2 covered the ten-image TACO pilot above plus the source-agnostic manual-upload annotation workflow; TACO import itself remains capped at exactly ten images and the full TACO import stays separately deferred. The master development order below is otherwise planned, not yet authorized.

  1. Build one excellent annotation screen.
  2. Manually test it on roughly 20–50 images.
  3. Fix the annotation UX based on that test.
  4. Import TACO.
  5. Begin the full manual review.
  6. Create Dataset v001 only after enough reviewed data exists.
  7. Build the M4 training agent and train Model v001.
  8. Only then connect inference back into the website.