52 Bukit Pasoh Road
#03-05
Singapore 089853

+65 6871 3405
[email protected]

Monday–Friday, 09:00–18:00 SGT. Closed on weekends and public holidays.

Bukit Pasoh Road, five minutes from Outram Park.

Journal

Field notes, 2026

Written by the studio, in the collective voice we use for the rest of the site.

Notes from the studio

These notes are from our first year. They are the arguments we keep having at the table on Bukit Pasoh Road: how to tell whether retrieval is actually working, how to collect an evaluation set that still looks like the questions staff type, how procurement tends to buy a demonstration, and how a system ages once the launch week is over. They are written for operators and for the people who have to keep the scores after we leave.

18 March 2026

Retrieval quality is a citation problem before it is a model problem

When a retrieval system fails, the first instinct is to swap the model. In the engagements we have run this year, the miss was more often a citation miss: the passage was in the corpus, the answer sounded right, and the file it pointed at was the wrong annex, the wrong year, or a slide that summarised a clause the contract had already replaced. Embeddings, the numerical summaries that let similar passages sit near each other, will happily retrieve a near neighbour. A near neighbour can still be the wrong clause for a partner to sign.

We now score three things on every evaluation question. Did the system refuse when the files had no answer. Did it cite a page a person agrees is the right source. Did the answer stay inside that page, or did it drift into fluent filler. The model matters for the third. The first two are design: chunk size, metadata, permissions, and the stubborn requirement that a reply without a citation is a miss even when a reviewer likes the prose.

A practical consequence is that we spend week two on the files. The prompt comes after the corpus can be trusted. Duplicate PDFs, scanned schedules without page numbers, and matter folders that mix drafts with executed copies will wreck a beautiful model. Cleaning is part of the retrieval product, with a size and an owner. If a folder cannot be trusted, we leave it out of the index and write that down, rather than hoping the embedding space will sort it out.

12 May 2026

How we collect an evaluation set without sanitising the questions

Quiet partitioned desks beside a window
The questions worth scoring are the ones typed at a quiet desk at the end of a shift.

An evaluation set is a list of questions drawn from real work, each paired with an answer a person in the firm has already judged. The set is useless if we polish it. Staff type with missing words, internal codes, and the name of a counterparty spelled three ways. If we rewrite those questions into brochure English, we are scoring a demonstration. We sit with the desk, ask for the last twenty things they typed to a colleague, and we paste them in as they are.

We also collect questions that should have no answer. A system that cannot say it does not know will invent a hold time, a clause, or a policy version. Those inventions are the expensive failures. We want them in the set from week one, labelled as expected refusals, so that a later model upgrade cannot buy fluency by guessing.

Ownership of the set has to sit with someone inside the firm. We can facilitate the first eighty questions. We cannot be the people who notice, in month five, that the night desk has invented a new shorthand. The handover includes a simple way to add a question, a schedule for reruns, and a rule that a model does not move into production until the set has been sat again. That sounds slow. It is slower than a rollback after a fluent wrong answer has already been acted on.

8 July 2026

Buying applied AI: a procurement note from the other side of the table

Procurement teams in Singapore are now being asked to buy “an AI”. The RFP often asks for a platform, a licence band, and a list of models. What the desk needed was a scored system on one corpus, with a human queue for the misses. We have sat through processes that selected the vendor who performed best on five polished prompts, then discovered in month two that the staff questions looked nothing like those prompts. The gap is not ill will. It is that a demonstration is easy to schedule and an evaluation set is work.

If we are invited to a procurement, we ask to replace the demo hour with a scoring hour. Give every shortlisted party the same thirty real questions, with answers a person has already judged, and a slice of the files those questions refer to. Compare citation quality, refusal quality, and the operational wrapping: logs, access, cost per query, what happens when the model is down. The wrapping is usually where the later pain lives, and it rarely appears on a slide of model names.

A second request: name the operator who will own the scores after go-live, in the contract, as a role. If that role is empty, you are buying a launch. We price discovery as a separate sprint so that a buyer can stop. Stopping after two weeks, with evidence that the files cannot support the task, is a procurement success. Paying for a year of a tool that cannot cite its sources is the more expensive kind of caution.

16 September 2026

Keeping a system alive after launch

Facade of an office building with blue window frames
The building stays. The corpus, the model, and the slang of the desk all move.

Launch week is the easy week. The evaluation set is still the one we collected together, the corpus is the one we indexed, and the people who trained on the tool still remember the refusal cases. Three months later a folder has been added, a model has been swapped because it was cheaper, and the night desk has a new code for damaged cartons. Drift is the name we give to that slow slide. Without a monthly score, drift looks like “the tool used to be better”, which is a feeling, and a feeling cannot be actioned.

Care, in our sense, is a short monthly loop. Rerun the set. Add the new slang. Read the journal of answers. Decide whether a model change is allowed. Train the next person on the roster. The dashboard is small on purpose: citation hits, expected refusals, unexpected refusals, and cost per query. If a number needs a meeting to explain, it is probably the wrong number.

The studios that disappear after go-live leave a system that only the original authors can nurse. We write the runbooks so that a new operator can run the monthly loop without a call to Bukit Pasoh Road. In our first year we have treated handover as part of hardening. The first monthly score is read together. After that, you can keep us on a thin retainer or you can run the loop yourselves. Either way the loop has to exist, or the system will quietly become a demo again.

The studio

If a note here describes your desk

Send a brief. We will tell you whether a discovery sprint is honest work for us, and what we would need in the first two weeks.