52 Bukit Pasoh Road
#03-05
Singapore 089853

+65 6871 3405
[email protected]

Monday–Friday, 09:00–18:00 SGT. Closed on weekends and public holidays.

Bukit Pasoh Road, five minutes from Outram Park.

Work

Four dossiers

All dates 2026

We write these as typical shapes of work. Names of organisations are withheld.

Anonymised case dossiers

These four write-ups are the centre of the site. Each one follows the same path: the situation we walked into, the system we built, what changed at the desk, and what we would do differently if we started again. They are written from engagements in our first year, founded in 2026 in Singapore, and they are anonymised at our clients’ request. The numbers describe scale and shape: how many files, how many weeks, how large a pilot group. They are not audited results, and they should not be read as a promise that your desk will move by the same amount.

Engagements are described in anonymised form at our clients' request. Figures indicate the scale and shape of the work rather than audited results.

Dossier

Professional services

Duration: 11 weeks

Stack: retrieval, pgvector, hosted model, eval set of 80

Associates were still hunting clauses by hand across several thousand PDFs.

Search across a working contract library

Desk with monitors and code beside a studio window
Prototype week: one corpus, one interface, scores on the wall.

The situation

A mid-sized professional services firm in Singapore kept contracts, schedules, and side letters across a matter system, a shared drive, and years of email. Associates spent the opening hour of many files searching for indemnity language, change-of-control clauses, and the last agreed version of a schedule. Partners still signed the advice. The pain was time, and the fear of quoting a superseded annex. Nobody trusted a single search box because previous attempts had answered fluently from the wrong year.

What we built

We indexed the working library and stored passages as embeddings, numerical summaries of meaning that let similar clauses sit near each other. Answers had to cite a file and a page. Before anyone beyond a pilot group typed a question, we sat with eight associates and collected eighty questions they had actually asked in the previous month, with answers a partner had already judged. The set included questions with no good source, so that invention could be scored as a miss. Access followed the same matter permissions as the human searcher.

What changed

After eleven weeks a pilot of twelve people used the system for first-pass search. The weekly note showed citation hits, honest gaps, and the questions that still failed. Associates still read the clause; partners still signed. The change we could see was where the first hour went: toward reading, away from hunting across three stores. We do not claim a percentage. We claim a desk that had a scored tool and a list of questions it was still allowed to miss.

What we would do differently

We would spend more of week one on the email PST files. The shared drive was cleaner than the mailboxes, and several of the missed questions lived only in a thread from 2024 that had never been saved to the matter. We would also add a “this looks like an older version” flag earlier, because version confusion was the miss that irritated partners most.

Dossier

Trade operations

Duration: 14 weeks

Stack: extraction, review queue, ledger write-back

A small desk retyping vessel names from scanned bills of lading.

Fields from trade documents into the ledger

People comparing printed charts and notes around a table
Agreeing which fields a clerk currently types by hand.

The situation

A regional trade operations team was keying invoice numbers, vessel names, and container counts from scanned bills of lading into the accounting system. Volume sat in the low thousands of documents a month. The desk was small, the scans were uneven, and a large vendor suite would have been a theatre around a problem that was still, at heart, typing. Wrong vessel names were noticed at month-end. The team’s own estimate of rework was “too much of Friday”.

What we built

We sampled a hundred documents and marked every field a clerk typed. The service extracted those fields, wrote confident rows into the ledger, and parked the rest in a review queue with the original image beside the guessed values. The confidence line was written down in week two, then adjusted once after the first fortnight of live queue. Stamps, shadows, and rotated pages were the usual reason a row waited for a person.

What changed

After fourteen weeks the job had a shape the supervisor could roster: a queue in the morning, exceptions through the day, a short report on how many rows posted without a touch. The silent wrong posting became rarer in the weekly sample we drew. We still treated every number on that report as a description of this desk, in this period, without an audit trail a board would want to publish.

What we would do differently

We would spend an extra week on the stamp-heavy scans before touching the clean digital PDFs. The clean PDFs made the prototype look further along than the queue would feel in week six. We would also involve finance in the confidence line earlier, because they were the ones who lived with a wrong posting, and operations were the ones who lived with a long queue.

Dossier

Warehouse operations

Duration: 9 weeks

Stack: Slack assistant, access mirroring, answer journal

The same dozen questions arriving after 21:00, answered by whoever was still awake.

A night-desk assistant in the tool the shift already uses

Conversation in a bright office corridor
Handover talk after a night-desk shadowing shift.

The situation

A warehouse and fulfilment operation in Singapore ran a thin night desk. Supervisors on the floor asked, in Slack, about hold times, packing exceptions, and which SOP applied to a damaged carton. The answers lived in a folder of procedures and in the memory of two people who were not always rostered. After 21:00 the same dozen questions returned. Previous chatbot trials had been retired because they answered from a brochure voice and could not point to a page.

What we built

We put a bounded assistant in the Slack workspace the shift already had open. It inherited the same folder permissions as the night supervisor. Every reply was written to a journal a daytime lead could read. The corpus was the SOP set plus a short list of exception codes. Questions outside that fence received a refusal and a named person to wake. We collected thirty-five real night questions as the evaluation set, including several that had no answer in the files on purpose.

What changed

Nine weeks later the night desk still needed a human. What changed was the first response to the repeating questions, and the daytime lead’s ability to see what had been said at 02:00. The journal showed refusals as well as answers, which is how we caught an SOP that the floor had already stopped following. We would not call that a headcount saving. We would call it a shift that had a cited first answer and a record.

What we would do differently

We would photograph the floor exceptions in week one. Several misses were questions about damage that the SOPs described in language the shift never used. We would also have the daytime lead join the evaluation scoring from the first rerun, because they were the ones who had to live with a wrong hold time.

Dossier

Claims operations

Duration: 12 weeks

Stack: queue ranking, reasons, human decision log

Reviewers picking the next file from a pile they could no longer see the bottom of.

Hints on a claims queue

Profile with binary digits projected across a face
Decision support stays a hint: a person still signs the file.

The situation

A claims operations group was working a queue that had grown faster than hiring. Reviewers chose the next file by habit and by whoever had chased them that morning. Simple files sat behind noisy ones. Supervisors wanted a ranking they could argue with, and a record of why a file had been pulled forward. They did not want a model that closed a claim. The licence to decide stayed with the reviewer.

What we built

We built a ranked list for the morning board: files the model guessed were short, files that looked like they needed a specialist, and files that had been waiting longest. Each row carried a short reason drawn from features a reviewer already used (age, missing documents, a keyword in the first page). The reviewer still opened the file. We logged the hint, the action, and whether the file returned to the queue. The evaluation set was a sample of past weeks that supervisors had already labelled as “should have been first” or “could wait”.

What changed

After twelve weeks the morning board had a shared order to disagree with. Supervisors used the log in the weekly huddle. Some hints were ignored on purpose, which we treated as useful signal rather than as failure of adoption. The queue did not vanish. The shape of the morning changed: fewer arguments about which pile to start, more arguments about whether a reason still made sense.

What we would do differently

We would have split “missing document” into three kinds in week two. The model treated a missing ID the same as a missing medical report, and reviewers did not. We would also have written the ignore-reason as a required field earlier, because the ignored hints taught us more than the accepted ones and we were slow to capture why.

Engagements are described in anonymised form at our clients' request. Figures indicate the scale and shape of the work rather than audited results.