How to read this case studyEvery number below is a target, not a trophy. The systems, the counters, and the arithmetic are written so a sceptic can check them.

Twelve hours, measured both ways.

A seven-week engagement at a 40-person B2B SaaS company: how Priyra's week was instrumented, what the five systems cost to build, and the 12.0 hours a week the ledger still reconciles to.

The engagement in one number

The 12-hour block.

Priyra's week held 26.2 hours of repetitive, low-judgment work. The engagement moved 12.0 of them into five systems — 46% of the baseline — and left 14.2 hours that genuinely need a person. The ledger at the bottom of this page reconciles to the same three numbers.

Before
26.2 h
Repetitive, low-judgment work, per week.
After
14.2 h
What still needs a person: exceptions, review, and judgement.
Recovered
12.0 h
46% of the baseline, on the same definitions.
Payback
2.1 weeks
$10,000 fee against $4,800 a week of recovered time.

Priyra, and the week she could not prove.

The problem was never that Priyra was slow. It was that her week was full of work that only needed her at the edges, and she had no number to take to the CEO.

Role
Chief Operating Officer
Company
40-person B2B SaaS company
Reports to
The CEO
Owns
Operations, finance ops, customer onboarding, weekly board reporting
Customers
About 140
Average contract
About $180k ARR

At 06:30 she was at the kitchen table with the shared inbox. Thirty messages from the day before, spread across a mailbox, a contact form, and a network inbox. Some were leads. Some were support. Some were vendors. She read each one, decided where it belonged, and forwarded it. By the time the house woke up, she had already done an hour of routing that needed a decent judgement on maybe three of the messages.

At 21:00 she was usually back at it, rebuilding the Monday pack because the numbers had moved since the morning, or chasing an approval that had sat in a thread all day. The board met on Tuesdays. The pack was due at 09:00.

The work was not difficult. It was constant, and it was hers because she was the only person with the context to unblock it. She estimated that twelve hours a week was recoverable, but she could not say where those hours were, and the CEO would not fund another hire on an estimate. The request that started the engagement was simple: prove the number first.

The cost was real, and larger than the fund-the-hire conversation. Her fully loaded time is the most expensive operator time in the company. Twenty-six hours a week of low-judgment work was roughly $10,480 a week, or just over $500k a year — more than the annual value of an average customer. That number is why the engagement got funded, and it is the reason the target was set at twelve hours rather than thirty.

The audit: two weeks of proof.

We did not build anything for the first four working days. We measured. The rule was that the baseline had to be one Priyra recognised and one the CEO could check.

The method

  • The audit ran across two calendar weeks: ten working days in total, logged to the day. Every task over ten minutes was logged with a start, a stop, and the system it belonged to. A weekly retrospective at the end of each week filled in what the day log missed.
  • A one-week instrumented re-measure, run in parallel with the second audit week. It used three sources that did not depend on Priyra's memory: calendar blocks tagged with the system name, sent-mail metadata (count and time, not content), and the ops ticket log's time field.
  • The diary found the tasks; the instrument sized them. Where the two disagreed by more than 30 minutes on a bucket, we took the instrument and wrote down why.
  • Everything was cut to the same five buckets used by the pattern taxonomy on the Work page, so a person reading both pages sees the same five words.

What it found

  • The baseline settled at 26.2 hours a week of repetitive, low-judgment work, not the 12 she had estimated. Her estimate was about the work she resented, not the work that cost her.
  • The instrumented week agreed with the diary to within about three hours a week. The diary was high on reading and low on the small interruptions between meetings, which cancels in the total and does not cancel in the cause.
  • The five buckets were lopsided. Two of them, the inbox and the market reading, were more than 45% of the total. Fixing the smaller three first would have produced a respectable number and missed the money.

Three surprises

  1. The inbox was five jobs wearing one name

    Reading, routing, replying, chasing, and remembering were tangled together. The largest single block inside the six hours was not answering messages. It was re-reading them to reconstruct context that the previous read had not recorded. A faster reply would have saved minutes; a routing layer that kept the context saved hours.

  2. The Monday pack was mostly not analysis

    Five and a half hours a week, and about three of them were copying figures between tabs and reconciling two sources that should have agreed. Priyra had been treating the whole block as the high-value part of her week. Most of it was joiner work.

  3. The reading was mostly repetition

    Six hours of market and competitor reading, and the instrumented week showed the same story arriving three times: once in a newsletter, once in an alert feed, once in a channel of shared links. Nobody was reading too slowly. They were reading the same thing more than once.

The thing she was sure was the problem

Priyra was certain the problem was meetings, and she arrived with a plan to cut a third of them.

The instrumented week showed about nine hours of meetings, most of them with a decision or a customer in them. The loss was not the meetings. It was the fifteen and twenty-minute gaps between them, where she reopened the inbox, re-read a message, found the pack had changed, and lost the thread before the next call. Cutting meetings would have created more of those gaps, not fewer. The engagement did not touch her calendar except to protect two blocks for the work the systems now hand back to her.

The five systems.

One per bucket, in the order they shipped. Each section follows the same five questions, then the numbers and the threshold it had to clear before it went live. The totals run to exactly 12.0 hours a week.

  1. 01

    Lead and inbound triage

    The inbox stopped being a routing table.

    The inbox pattern on the Work page
    What it looked like before
    Thirty messages a day arrived in three places: a shared ops@ mailbox, the website contact form, and a network inbox nobody owned. Two operations staff read every one and decided whether it was a lead, a support question, a billing problem, a partner note, or spam. Priyra took the ambiguous ones and every message from an account that looked large. Replies went out when someone reached that part of the queue, which on a busy Tuesday could be the next morning. An enterprise prospect got the same two-day response as a password reset.
    The broken mechanism
    The routing decision needed account context, not keywords. The same sentence — can we add three more seats? — is a support task from an existing customer and a sales opportunity from a prospect. A rules-only filter misrouted both ways, and a checklist could not hold the rule because the rule depended on the account, not the words. Nobody could say how many messages had been misrouted, because nobody was counting. That is the difference between a busy inbox and a broken mechanism: the failure was invisible until the counter existed.
    What we built
    A watcher on the mailbox, the contact-form webhook, and a scheduled export from the network inbox push each message into one queue table. A preprocessing step strips signatures and quoted replies, normalises the sender domain, and joins the message against the CRM and the billing system to attach four facts: lifecycle stage, ARR band, open tickets, and account owner. A model call classifies the message into one of five routes and returns a constrained object — route, confidence, reason, suggested owner — with no free text. A deterministic layer above the model handles the cases that must never be guessed: known vendor domains, messages tied to an open ticket, and any account above $150k ARR. Routing writes a task in the ticketing system or a deal task in the CRM, sets an owner, and drafts a first reply for the two lowest-stakes routes.
    Guardrails and escalation
    The system never sends mail. Drafts are created as drafts; a person sends them. It refuses to decide refunds, contract terms, security questionnaires, and anything from a customer flagged at risk — those go straight to a person with the context attached. Any classification below 0.8 confidence, and every enterprise account, lands in an exception queue reviewed twice a day. The queue raises an alert if it holds more than six items or if anything has waited four business hours. Every message row keeps the raw headers, the model version, the classification, the confidence, the rule that fired, and any human override; the override becomes a labelled example in the next week's test set.
    How the saving was measured
    The queue table records received_at and first_human_action_at. The weekly view counts the messages a person had to read and decide, and sums the minutes recorded against each. In the baseline, that was every message: about 6.0 hours a week across two people plus the messages Priyra handled. After routing went live, the counter only advanced for the exception queue, the enterprise route, and any message a person pulled back. The four-business-hour response target is measured from the same timestamps, so speed and volume share one instrument rather than two opinions.
    Before
    6.0 h
    After
    3.0 h
    Saved
    3.0 h

    To go live

    Routing had to reach 92% accuracy on a 200-message set labelled by the team, with zero enterprise leads auto-routed to a low-priority queue, and a median first human action under four business hours. It had to beat the manual baseline on misroutes, not only on time. It cleared the accuracy bar at 94% and the enterprise rule at 100%.

  2. 02

    Weekly reporting pack

    The Monday pack assembled itself before she opened the laptop.

    The reports pattern on the Work page
    What it looked like before
    Every Monday, Priyra and a finance analyst built the pack that went to the board and the leadership team: ARR and MRR movement, churn, pipeline, support volumes, onboarding status, and a page of commentary. The numbers came from Stripe, the CRM, the product database, the support tool, and a spreadsheet of onboarding milestones. It took about five and a half hours across two calendars. Most of that was copying figures between tabs, checking that Stripe and the product database agreed, and rewriting the commentary on Monday morning because the numbers had moved overnight. The Tuesday board meeting set the deadline, so the pack was always built under it.
    The broken mechanism
    The pack mixed two jobs. Assembling the numbers is mechanical; explaining them is a judgement call. The mechanical part failed because each source had its own definition, timezone, and lag, and a template could not decide which number was authoritative when two disagreed. The commentary was rewritten every week because it was written against numbers that were still moving. A better checklist would have made the disagreement more organised, not smaller. The failure was the absence of one agreed definition per metric, not the absence of effort.
    What we built
    A nightly job materialises one table with a single agreed definition per metric: ARR as the sum of active subscription MRR, churn on a 30-day trailing window, pipeline by stage, and so on. On Monday at 05:00 a job pulls the week's movement, compares it against the prior week and the same week a year ago, and flags anomalies such as churn above 1.5 times the trailing eight-week median. It writes one structured input row. A model call drafts the commentary from that row only. Every figure in the prose is a substitution from the structured fields; the model cannot introduce a number. The output is a draft in the house format plus a Markdown copy for the archive.
    Guardrails and escalation
    A validation gate reads the draft, extracts every numeric token, and fails the run if any figure is not in the structured input. The gate also fails if a metric is missing or an anomaly is unflagged. The model is not allowed to smooth, project, or round; it can only reference what the job computed. The draft is never sent. Priyra reviews it on Monday morning and edits the narrative herself. Each run stores the input snapshot, the model version, the validation result, and a diff of her edits. A gate that is always bypassed is not a gate, so the pass rate is logged weekly.
    How the saving was measured
    The warehouse records the run start and finish and the hash of the input snapshot. Human time is counted from two places: the one-hour calendar block Priyra keeps for review, and the analyst's time against the reporting ticket in the ops log. Before, the two calendars showed 5.5 hours a week on the pack. After, the counter covered the review hour, the analyst's exception checks, and the occasional definition fix: 3.1 hours. The validation gate's pass rate and the count of editorial edits are logged separately, because a draft that needs a full rewrite has saved formatting time and not judgement time.
    Before
    5.5 h
    After
    3.1 h
    Saved
    2.4 h

    To go live

    The pack had to match the hand-built pack for every figure on two consecutive Mondays, with the validation gate passing, before the analyst stopped building it by hand. It also had to detect one seeded anomaly the team planted in the data. It passed the figure check and the seeded anomaly, and it failed the first gate run because a footnote number was not in the structured input — which is exactly what the gate is for.

  3. 03

    Approvals and document generation

    Approvals stopped waiting on whoever noticed.

    The ops pattern on the Work page
    What it looked like before
    Requests for spend, vendor onboarding, contract review, and new-hire setup arrived by email or in a message thread, in no particular format. Priyra approved most of them, but only after asking for the missing information that should have been there the first time. The same three documents — an SOW, a vendor security summary, and an onboarding checklist — were rebuilt from scratch every time. Nobody could say what was waiting on whom. The work was small and constant: four hours a week, rarely in one block, always interrupting something else.
    The broken mechanism
    The bottleneck was not the decision. It was that the request arrived without the facts the decision needed, and the chase for those facts was manual. A checklist sent to the requester did not fix it, because people filled in the checklist after the fact and the approver still had to verify each line. The documents were rebuilt because the source data lived in three systems and a person was the join. Automating the approval click would have moved the same chase around; the fix had to be the schema and the joins.
    What we built
    One request form, a small internal page, with a fixed schema per request type. Submitting it creates a row with a status. A rules layer checks the request against policy: spend under a threshold in the requester's budget, a vendor already approved, a standard contract template. If everything passes, the system generates the document from the request row plus the CRM and finance records, files it in the right folder, and moves the status to approved. If a fact is missing, the system asks the requester for that one fact and re-enters the queue; it does not send the request to Priyra until the schema is complete. A separate queue holds anything the rules cannot clear.
    Guardrails and escalation
    The system refuses to approve anything that touches pricing, headcount, legal terms, or a new vendor it does not already know. It cannot sign, send externally, or change a contract clause. Any request that mentions a regulated data category or a non-standard liability term is routed to a person and flagged. The audit trail stores the request row, the policy checks with their inputs, the generated document hash, the approver, and the time from submission to decision. Documents are versioned; the generated file is never overwritten. Human decisions stay with every exception, every request above the spend threshold, and every contract that is not the standard template.
    How the saving was measured
    The request table has submitted_at, ready_at, and decided_at. The weekly counter sums the minutes between submission and ready, and between ready and decided, per request type, and counts how many requests needed a chase. In the baseline, the chase and the rebuilding were the four hours. After, the counter only advanced for exceptions and for the standard decisions Priyra still made. The time from submission to a complete request fell from about two days to under four hours, which is the number the finance team actually felt. Counts come from the table, not from anyone's recollection.
    Before
    4.0 h
    After
    2.4 h
    Saved
    1.6 h

    To go live

    It had to clear 90% of standard requests without a person touching them and produce documents that matched the previous hand-built versions field for field on a 30-document sample. It also had to bring median submission-to-ready below four hours. The first build passed the document check and failed the chase counter, because requesters emailed Priyra directly instead of using the form. The form went into the mailbox auto-reply in week five, and the counter moved.

  4. 04

    Market and competitor digest

    The reading nobody had time for, triaged and cited.

    The research pattern on the Work page
    What it looked like before
    Priyra read for the job: a dozen industry newsletters, three competitor alert feeds, a channel of shared links, and the analyst notes the sales team forwarded. She read to stay credible in board and customer conversations, not because anyone asked her to. It took six hours a week and much of it was the same story in three places. When she travelled, the reading did not happen, and by the next week the pile was unreadable. The reading list was also the first thing to be cut when something urgent arrived, which meant the most valuable part of it arrived late.
    The broken mechanism
    The reading was not one task. It was collection, de-duplication, triage, and interpretation tangled in the same hour. The cost was not only reading time; it was that the reading happened late, so the useful signal reached a customer conversation after the fact. A faster reader would only have read the duplicates faster. The mechanism needed a stage that removed repeats before a person saw them, and a source of record so an item could be reopened without re-reading the pile.
    What we built
    A scheduled collector pulls an allowlisted set of sources: newsletters by mailbox rule, competitor feeds by RSS, and the saved-links channel. Each item is reduced to its text, de-duplicated against the last 60 days by near-duplicate similarity, and tagged by topic and source reliability. A model call writes a two-sentence summary per item and extracts the claims with a source link attached. A ranking step orders items by relevance to a fixed set of business themes — pricing changes, product launches, funding, regulation, hiring signals — and writes one page with a short list of items that may need a response. The digest is not a report; it is a reading list with the repeats removed and the yes/no left to her.
    Guardrails and escalation
    The system cites a source for every claim and refuses to state an implication or a recommendation. It does not summarise paywalled content it cannot access; it flags the headline only. Anything below the relevance threshold is dropped, and anything the classifier marks as a rumour or unverified is labelled as such. The digest is capped at one page so it cannot become another pile. Every item keeps its source URL, the similarity cluster it belonged to, and the score it received, so a dropped item can be explained rather than mourned.
    How the saving was measured
    The collector logs items in, items deduplicated out, and items kept. The weekly counter measures the hours Priyra spends reading the digest and the hours she spends reading raw sources, separately, because the whole point was to move time from raw reading to the digest. In the baseline, both were the same six hours. After, the digest read was 2.8 hours and raw reading was near zero unless she chose to open a source. The duplicate-removal rate is logged per week and settled between 40% and 55%, which is the finding the audit predicted.
    Before
    6.0 h
    After
    2.8 h
    Saved
    3.2 h

    To go live

    It had to drop duplicates without dropping a competitor pricing change, keep the digest to one page, and hold a source link on every claim. On a two-week trial it missed one relevant item the team had seen, which was enough to keep a weekly human sweep of the source list for the first month. After the sweep found nothing for three weeks, it was retired.

  5. 05

    Invoice reconciliation

    Month-end matching down to genuine exceptions.

    The reconciliation pattern on the Work page
    What it looked like before
    Every month end, a bookkeeper and Priyra's operations lead matched purchase orders to invoices to payments line by line. The average week was 4.7 hours, most of it in the last three days of the month. The work was checking three lists against each other and writing down the differences. Two errors reached payment in the previous year, both caught late by the vendor rather than by the process. The whole task waited on two people's calendars, so a holiday week pushed it into the next month.
    The broken mechanism
    The matching rule was simple in the common case and impossible to write down in the edge cases. Partial payments, credits, currency, and multi-line invoices each needed a different comparison, and the rule depended on the vendor's own history. A template could not hold it, and a person could not apply it consistently under time pressure at month end. The reconciliation also had no memory: every month the same vendor's same quirk was rediscovered. The mechanism failed on memory and consistency, not on diligence.
    What we built
    Invoices arrive in a dedicated mailbox or through the finance system. A parser extracts header fields and line items with a confidence per field. A matching engine compares against open purchase orders and recorded payments under rules that are specific per vendor: tolerance band, currency handling, partial-payment logic, and credit-note treatment. Exact matches are written to the ledger and marked reconciled. Everything else goes to an exception queue with the evidence assembled: the three documents, the fields that differ, and the vendor's prior handling. A model call drafts a plain-language reason for each exception, never a decision. Human decisions stay with every mismatch, every payment over a fixed threshold, and every first-time vendor.
    Guardrails and escalation
    The system never releases a payment. It writes a reconciliation record and a status; a person approves the payment run. It refuses to match across currencies without a same-day rate, refuses to net a credit against an unrelated invoice, and refuses to touch an invoice flagged as disputed or under legal review. The audit trail stores the parser confidence, the rule-set version, the three source documents, the match decision, and the approver. A mismatch approved with an override is recorded and becomes a rule candidate, so the same quirk is answered by the system the second time.
    How the saving was measured
    The reconciliation view has invoice_received_at, matched_at, and approved_at. The weekly counter sums the bookkeeper's and operations lead's logged minutes on reconciliation tasks and counts the exceptions a person had to open. In the baseline, the counter showed 4.7 hours. After, it showed 2.9 hours, nearly all of it in the exception queue. The count that moved most is the exception rate: from treating every line as an exception to roughly one invoice in six, and falling as vendor rules were added.
    Before
    4.7 h
    After
    2.9 h
    Saved
    1.8 h

    To go live

    It had to match to the penny on a sample of three months of historical invoices with no false match, and hold every genuine mismatch for a person. A false match is worse than a missed one, so the bar was zero. It passed the historical sample after two rebuilds of the currency rule, and the first live month produced eleven exceptions against an expected nine.

The ledger.

The five systems, the hours each one moved, and the residue that stays with a person. The arithmetic is the whole argument, so it is left where it can be read: before minus after equals saved on every row.

Hours a week before and after, by system. All figures are hours a week.
No.SystemBeforeAfterSaved
1Lead and inbound triageinbox6.03.03.0
2Weekly reporting packreports5.53.12.4
3Approvals and document generationops4.02.41.6
4Market and competitor digestresearch6.02.83.2
5Invoice reconciliationreconciliation4.72.91.8
Total26.214.212.0

Before 26.2 − after 14.2 = 12.0 hours a week recovered, 46% of the baseline. All figures are hours a week.

What we deliberately left alone.

Not everything on the list should be automated, and one of the honest parts of the engagement was the list of things we declined to touch.

  • Pricing exceptions. Anything other than the published discount band is routed with the account history attached, but Priyra decides. A discount is a relationship and a margin decision at once, and a wrong one is expensive in both directions.
  • Hiring. The hiring pipeline was untouched. Scheduling interviews and collecting feedback is mechanical, but a faster loop was not the problem, and the judgement in a hire is the whole point.
  • Legal, security, and contract terms. Liability, data-processing questions, security questionnaires, and non-standard clauses. The system assembles the facts and the prior answers; a person answers each one and signs it.
  • Board and investor judgement. The pack is assembled automatically. The story about the quarter is not. The model is barred from drawing conclusions the data does not contain, and Priyra writes the narrative.
  • Performance and personnel decisions. No part of a person's review, workload, or pay was automated. It was out of scope, and we said so at the start rather than discovering it later.
  • First-time vendors. Every new vendor is reviewed by a person with the documents assembled, because the trust decision has no history to match against. The system makes the review fast; it does not make the decision.

The week after.

Twelve recovered hours do not stay recovered on their own. A released hour goes back to whatever shouts loudest unless it is named and protected, so the allocation is part of the engagement, not a footnote to it.

Where the twelve hours went

  • 3.0 h

    Customer and partner conversations

    A standing call with the ten largest accounts each quarter, and a short call whenever an account goes quiet. This was the first thing to disappear when the week got busy, and it was quietly costing renewals.

  • 2.0 h

    Coaching and one-to-ones

    A weekly thirty-minute one-to-one with each of the four operations staff. She had been cancelling these first, which is how a team loses context.

  • 2.0 h

    The process backlog

    Two hours a week on the list of internal fixes that never reached the top. The onboarding handoff between sales and operations was first, and it removed a recurring weekly irritation.

  • 2.0 h

    Finance and board depth

    Scenario modelling instead of just reporting. Two or three forward cases for the board rather than a single number defended in the meeting.

  • 1.0 h

    Friday exception review

    A protected hour to clear the week's queues, read the override log, and decide which rule to change next. The systems get better only if someone reviews the exceptions.

  • 2.0 h

    Protected thinking time

    Two mornings, 07:00 to 08:00, blocked and defended. Unglamorous and easy to steal, which is why it is written here as an allocation rather than an intention.

A day now

  • 06:30 — 07:00

    Reads the one-page digest. Ten minutes if the week is quiet, half an hour if a competitor moved. The shared inbox is no longer the first thing she opens.

  • 07:00 — 08:00

    Protected thinking time on Tuesdays and Thursdays. Blocked, defended, no meetings and no inbox. This is the block that would have been the first casualty of a busy week before.

  • 08:00 — 09:00 (Mondays)

    Reviews the generated pack, checks the anomalies, and writes the narrative. The assembly happened at 05:00 without her.

  • 09:00 and 16:00

    The exception queues, twice a day as a habit rather than an event. She approves, corrects, or signals a change. The corrections become test cases, so the queue should shrink over time.

  • 10:00 onward

    Meetings, customer calls, and the work only she can do. The routing decisions happen in the queue, not in her head.

  • 21:00

    No longer a second working block. If something is waiting, it has an owner and a timestamp in a queue, not a place in her memory.

The 14.2 hours that remain are not waste. They are the exception queues, the reviews, and the judgement the systems deliberately hand back. Some of the residue is a choice.

Seven weeks, one re-do.

The rollout ran in week order, smallest blast radius first. One week failed its own threshold and was withdrawn and rebuilt. This is the version with the failure left in, because the failure is the useful part.

  1. Week 1

    Instrument and baseline

    Shipped
    The task timer, the mailbox metadata export, the ticket-log query, and the skeleton warehouse table. The five buckets were named before any build started.
    Broke
    The first two days of self-logging were written from memory at 18:00 and were unusable. We switched to a timer with two-minute granularity and threw the two days away.
    Outcome
    Baseline fixed at 26.2 hours a week. The warehouse named the metrics the reporting system would later have to satisfy.
  2. Week 2

    Inbox, first pass

    Shipped
    The message queue, preprocessing, and the classifier for the two clear routes, support and sales.
    Broke
    The classifier read quoted reply text and misrouted whole chains. Stripping quotes and classifying only the thread tail fixed most of it, and exposed the rest as account-context problems, not model problems.
    Outcome
    Routing accuracy 84%, below the 92% bar. The enterprise rule and the confidence gate came next rather than shipping a weak classifier.
  3. Week 3

    Inbox, second pass, and the exception queue

    Shipped
    Deterministic enterprise and vendor rules, the confidence gate, the exception queue, the audit trail, and the first 200-message test set.
    Broke
    Alerts fired on every exception and the team muted the channel within a day. The queue moved to a twice-daily digest, with an alert only above six items.
    Outcome
    94% accuracy. Inbox shipped with a human still sending every reply.
  4. Week 4

    Reporting, and the week that had to be re-done

    Shipped
    The metric definitions, the nightly table, the draft generator, and the validation gate.
    Broke
    Exactly what the gate exists for. Stripe and the product database disagreed at the month boundary because the product database lagged by a day, and the first gate run failed. The pack did not generate.
    Outcome
    The week's deliverable was withdrawn. Two metrics were redefined with a documented lag, and the build was redone the next week. Shipping the pack on the old definitions would have put a wrong number in front of the board.
  5. Week 5

    Reporting, rebuilt, and approvals

    Shipped
    The corrected metric definitions with the documented lag, the anomaly check, the reporting build, and the approvals form with its schema and rules layer.
    Broke
    The document generator produced the right fields in the wrong order on the standard SOW. A template fixed it without touching the model.
    Outcome
    The pack matched the hand-built pack on two consecutive Mondays. Approvals cleared 91% of standard requests untouched, with the chase counter still flat because people were emailing instead of using the form.
  6. Week 6

    Market digest

    Shipped
    The collector, de-duplication, topic ranking, the one-page digest, and source links on every claim.
    Broke
    The first de-duplication pass dropped a competitor pricing change as a near-duplicate of an older post. The similarity threshold was tightened and the source allowlist made explicit so a whole source could never be silently dropped.
    Outcome
    Between 40% and 55% of items removed as repeats. The weekly human sweep stayed for a month and was then retired.
  7. Week 7

    Reconciliation and handover

    Shipped
    The invoice parser, the per-vendor rules, the matching engine, the exception queue, the runbook, the evaluation sets, and the operations handover session.
    Broke
    The currency tolerance matched an FX movement as a full match on the first live run. Per-currency tolerances and a same-day-rate requirement fixed it before the payment run.
    Outcome
    Full rollout. The five systems ran in production, and the ledger was measured over the following two weeks on the same counters as the baseline.

How the number is counted.

A time-saving claim is only as good as its counter. This is the method, the way to reproduce it, and the way to falsify it.

The counting method

  • Every system writes its own counter: the message queue, the report run, the request table, the digest collector, and the reconciliation view. Nobody reports hours from memory after the baseline.
  • Human minutes come from two sources that do not depend on the model: calendar blocks tagged with the system name, and the ops ticket log's time field. Where the two disagree, the log wins.
  • The baseline used the same two sources, reconstructed, plus the two-week diary. The diary found the tasks; the instrument sized them.
  • The weekly view sums by bucket and reports before and after on the same definitions, for the same person and the same month, not across teams or quarters.

How a sceptic would reproduce it

  • The runbook holds the SQL for the weekly view, the raw event tables, and the 200-message routing test set. Anyone with read access can rerun the view and get the same weekly totals.
  • The validation gate can be re-run against the same input snapshots, so the reporting figures can be checked independently of the prose.
  • The diary is kept as an appendix: 14 days of task entries, which is the rawest artifact in the engagement and the easiest to argue with.
  • Ninety days after rollout the same two-week instrumented measure is repeated, by someone who did not build the systems, and compared to the baseline in this page.

What would falsify this claim

The claim would be falsified if a fresh two-week instrumented measure, run by someone who did not build the systems, showed the five buckets falling by less than about nine hours a week, or if the exception rate rose over time, which would mean the systems were handing work back rather than absorbing it. It would also be falsified if the calendar still showed the 06:30 and 21:00 blocks after 90 days, or if the validation gate's pass rate fell, which would mean the reports were being corrected by hand each week. In this composite those are the tests we set, not results we are asserting. The numbers are the target arithmetic; the method is what makes the target checkable.

Honest limitations

  • Novelty effect: productivity usually looks best immediately after a change. That is why the target is re-measured at 90 days, not on the first week.
  • Self-report risk: the baseline diary is the weakest instrument. The instrumented week exists to correct it, and the two disagreed by about three hours a week.
  • One company, one person: the arithmetic belongs to Priyra. The split between the five systems will differ anywhere else, and the total may not move at all.
  • Buckets are not airtight: a message about a contract is partly inbox and partly ops. We counted it where the first human action happened and left it there for the whole engagement.
  • The target is capacity, not cash: twelve hours is released time. Whether it becomes revenue depends on what the person does with it, which is why the week-after allocation is part of the account.

The fee, the payback, and the run cost.

The payback number is only meaningful if the fee, the value of an hour, and the ongoing cost are all stated. Here they are, together, so the division can be checked.

Engagement fee
$10,000
A seven-week build across five systems.
Loaded cost of the executive's time
$400 / hour
Covers salary, bonus, employer costs, and overhead.
Recovered time
$4,800 / week
12.0 recovered hours × $400, on the same loaded rate.
Payback
2.1 weeks
$10,000 ÷ $4,800 = 2.08 weeks, rounded to one decimal.
Run cost
$210 / month
Model calls, infrastructure, and monitoring at the volumes above, about $2,520 a year.

The run cost is stated beside the fee on purpose, because it is real and it belongs in the picture. The payback above is the fee against recovered time only: 12.0 × $400 = $4,800 a week; $10,000 ÷ $4,800 = 2.08 weeks, rounded to 2.1. It does not include the $210 a month run cost, which sits beside it as the ongoing side of the same decision. At the loaded rate the recovered time is worth about $230,000 a year over 48 working weeks, which is the kind of number that sounds too good, so the honest version is the smaller one. The engagement bought back capacity that avoided a hire. It did not put $230,000 in the bank. The payback is the number that survived the scepticism.

What we would do differently.

The engagement worked, and several parts of it worked by accident. These are the parts we would change on the next one, including the parts that did not work.

  • Instrument before we build

    We spent the first four days building the warehouse before the counters were agreed and rewrote two views. The measurement should be first, and it should be boring.

  • The definitions were the hard part

    The model work took days; the argument over what ARR means took a week. It should have been a week-one deliverable with the finance owner in the room, not a week-four surprise.

  • We over-built the exception queue

    The first version was an internal web app with filters and assignments. It went unused. A shared channel and a spreadsheet did the job, and the queue is still a channel.

  • We underestimated habit

    Two people kept the old spreadsheet for a month while the new system ran beside it. That month cost more than any technical fix. Next time the old path is switched off on a named day.

  • We let the model write the implications

    The first market digest ended each item with a so-what. It was generic and often wrong. We removed it. The digest now summarises and cites, and Priyra writes the implication. The system got better when it did less.

  • The joins, not the model

    Most of the failures were domain matching, timezone lag, and duplicate detection. No model choice would have fixed a message table that could not join to a customer. The unglamorous data work was the actual project.

Which week is yours?

Tell us where the time goes. We will tell you what the first system would be, and what the twelve hours would cost.