Babylon 2k
Proposed learning design

Make every interview answer earn its place.

Collect an answer when it changes a useful decision, supplies evidence or makes Beth easier to use. Name the feature that consumes it and the test that proves its value.

This report audits the 24 existing interview topics. It proposes the missing learning and approval process; it does not change your saved answers, current quotations or live Beth behavior.

Start with the founder’s outcome.

Beth needs to understand the concern, ask the next useful question, explain the right work, calculate a bounded fee and recommend someone eligible to deliver it.

The AI already has general language skills. Your team supplies the local service definitions, private costs, exceptions, professional judgement and examples that make those skills useful for Babylon.

Before asking a stakeholder

  1. What decision changes if this answer differs?
  2. Which component reads the structured value?
  3. Is it already available from an authoritative source or directory?
  4. Who can verify it, with what evidence?
  5. What test proves its usefulness or identifies a failure?
Founder concern → useful question → confirmed or explicitly uncertain fact → approved service rule → bounded scope → calculated fee → eligible specialist → measured job outcome

Six destinations for the answers.

The first version should use approved records, executable rules and an existing language model. Raw interviews are not automatically model-training data.

DestinationWhat the answer changesWhat consumes it
Cited knowledgeRetrieve the approved, applicable source when answering.Guidance retrieval and answer grounding
Questions and service rulesInterpret founder replies, ask the next useful question and choose bounded work.Profiling, scope and dependency evaluator
Pricing configurationCalculate effort, cost and fee from approved numeric units and policy.Deterministic pricing calculator
Specialist eligibility and matchingFilter by qualification, capability, conflict and location before ranking.Live roster and matching engine
Delivery and case workflowSet document stages, deliverables, acceptance and accountable handoff.Case and task system
Examples and quality testsShow good conversation behavior and compare outcomes against unseen expert-labelled cases.Approved prompt examples and independent test runner

Prices, capacity and credentials stay in updateable records. The quote calculator performs arithmetic; the language model explains approved tool results. Stable conversation examples may guide the model through instructions before any model customization is considered.

Example: one question changes the quote.

Specialist’s answer: “Our monthly base covers 100 agreed accounting entries, clean records and two accounts. Each additional started block of 100 entries is extra work.”

Structured rule: a numeric allowance of 100, an increment of 100, one entity and one month, linked to BK01 and BK02.

Beth’s founder question: “About how many sales, purchase and expense entries do you handle each month? A rough count is fine; we will confirm the agreed work unit.”

Founder’s answerExpected behavior
80 agreed entries, clean records, two accountsBK01 × 1; BK02 × 0; agreed monthly period stated.
280 agreed entries, clean records, two accountsBK01 × 1; BK02 × 2; numeric calculator uses approved hours, role costs and policy.
Not sureentrycount remains unknown; ask one useful clarification or show a clearly bounded, reviewable estimate. Do not infer a count from revenue.
Records several months behindAssess a separately scoped cleanup phase; do not silently include unlimited historical reconstruction in a monthly fee.

Illustrative future contract using existing pilot service IDs; not an approved Babylon rule. The current engine hardcodes some allowances, so it must be connected to structured rules before edited definitions can reliably change scope. Applicable tax or other services are separate decisions.

Review and publish a complete version.

  1. Captured

    Keep the original answer and respondent attribution. Unknowns stay unknown.

  2. Structured draft

    Extract typed candidates. Reject missing units, periods, source scope and invented defaults.

  3. Needs resolution

    Conflicting definitions, currencies or rules remain unpublished; compare like segments before resolving.

  4. Reviewed

    Service owner reviews scope, finance reviews costs, expert reviews applicability and citations.

  5. Approved candidate

    Named owners sign off. Keep estimated and measured inputs visibly distinct.

  6. Tested release

    Run the reference cases and critical checks. Persist the full version bundle.

  7. Published

    Beth uses only this approved version. Keep the previous version available for rollback.

  8. Review due / superseded

    Changed source, credential, policy or measured outcome creates a new draft; never overwrite history silently.

Who decides?

Service owners approve units, scope and deliverables. Finance approves cost and commercial policy. Experts approve applicability, citations and advice boundaries. Operations owns capacity and handoff.

What stays separate?

Keep original answers, evidence, extracted candidates and approval decisions separately. Identity and private payout data belong in controlled records. The founder conversation receives only relevant approved guidance and scoped tool results.

Proposed record types and relationships
InterviewSession

id, respondent_id, verified_role, interviewer_id, service_segment_scope, started_at, submitted_at, answer_revisions, response_status

CandidateInput

id, input_type, source_session_id, source_question_id, exact_answer_reference, typed_value, unit, period, entity_segment, evidence_ids, unknown_fields, conflict_group, valid_from, valid_until, proposed_by, review_status

Evidence

id, kind, source_reference, authority, clause, effective_date, checked_date, anonymised_job_id, estimate_or_measured, superseded_by

ReviewDecision

candidate_id, reviewer_id, verified_role, decision, reason, reviewed_at, resolved_conflicts

BethRelease

id, service_dictionary_version, package_version, rule_version, pricing_version, source_version, prompt_example_version, approved_by, published_at, previous_release_id

ReferenceScenario

id, underlying_case_group, journey, input, expected_facts, expected_clarification, expected_service_scope, reference_fee_or_bound, citation_requirements, required_escalation, prohibited_claims, expert_labels, development_or_holdout, synthetic_or_actual

TestResult

scenario_id, release_id, model_and_prompt_version, run_id, actual_facts, actual_scope, actual_quote, question_count, citation_support, critical_failures, human_review, paired_comparison

JobOutcome

anonymised_job_id, release_id, service_entity_period, planned_hours, actual_hours_by_role, realised_cost, scope_changes, quality_result, price_override_reason, calibration_review

Each approved artifact links back to the interview session, question, answer and evidence. A release pins the service, rule, pricing, source and example versions; reverting a release restores the previous complete bundle.

Audit all 24 interview topics.

Prefill known directory and catalogue details. Ask stakeholders to confirm or resolve them. Capture rules and numbers in structured fields, and obtain dynamic availability and outcome evidence from operations.

Prove the answers help before expanding.

Start with three journeys: everyday monthly records, catch-up / missing records and business setup scope.

Suggested starting panel: three delivery specialists, one finance/commercial owner and one expert. Planning numbers only; they do not establish statistical sufficiency.

Build 45 expert-labelled cases

For each journey, use typical, more complex, uncertain, conflicting-information and boundary cases. Draft three distinct cases per category. Use 30 for development and keep 15 unseen for the comparison.

Split by underlying case, keeping paraphrases together. Mark synthetic cases. Reference answers, required facts and acceptable scope are approved by humans; an AI-generated answer is not automatically the correct label.

Compare three versions

A: current Beth.
B: the same model with approved interview-derived records, rules and examples.
C: B with one input family removed at a time.

Keep the model, prompt budget, tools and source snapshot consistent. Repeat the held-out cases three times and blind the reviewer to the version.

Measure useful outcomes

Proposed release checks: no critical failures in the suite; start with a 90% expert-approved scope target and no regression in unnecessary services or question burden. Stakeholders must agree targets before running the study.

Passing a small suite is a release check, not a guarantee of production accuracy. No improvement scores are claimed in this report. Compare pricing to reviewed calculations first and actual job costs later. Sales conversion alone cannot establish correctness.

Keep an input when it improves a measured task or prevents a material error. Keep critical eligibility and advice-boundary checks even when their average benefit is small. Prefill, simplify or defer inputs without a relevant consumer.

Learning after real jobs: record planned versus actual hours by role, realised cost, record quality, scope changes and delivery outcomes. The team reviews whether the rule or estimate was wrong before publishing a new version.

Model training is a later, evidence-based choice.

Use approved knowledge retrieval, business rules, quote tools and good examples first. If a persistent behavior failure remains, evaluate model customization against the same unseen cases. Dynamic legal guidance, partner costs and availability should remain maintainable records.

The OpenAI fine-tuning documentation currently says the platform is winding down and unavailable to new users. This design therefore depends on neither new fine-tuning jobs nor the hosted Evals platform; the reference cases and comparisons can run independently.

Official documentation informing the approach