Make every interview answer earn its place.
Collect an answer when it changes a useful decision, supplies evidence or makes Beth easier to use. Name the feature that consumes it and the test that proves its value.
This report audits the 24 existing interview topics. It proposes the missing learning and approval process; it does not change your saved answers, current quotations or live Beth behavior.
Start with the founder’s outcome.
Beth needs to understand the concern, ask the next useful question, explain the right work, calculate a bounded fee and recommend someone eligible to deliver it.
The AI already has general language skills. Your team supplies the local service definitions, private costs, exceptions, professional judgement and examples that make those skills useful for Babylon.
Before asking a stakeholder
- What decision changes if this answer differs?
- Which component reads the structured value?
- Is it already available from an authoritative source or directory?
- Who can verify it, with what evidence?
- What test proves its usefulness or identifies a failure?
Six destinations for the answers.
The first version should use approved records, executable rules and an existing language model. Raw interviews are not automatically model-training data.
| Destination | What the answer changes | What consumes it |
|---|---|---|
| Cited knowledge | Retrieve the approved, applicable source when answering. | Guidance retrieval and answer grounding |
| Questions and service rules | Interpret founder replies, ask the next useful question and choose bounded work. | Profiling, scope and dependency evaluator |
| Pricing configuration | Calculate effort, cost and fee from approved numeric units and policy. | Deterministic pricing calculator |
| Specialist eligibility and matching | Filter by qualification, capability, conflict and location before ranking. | Live roster and matching engine |
| Delivery and case workflow | Set document stages, deliverables, acceptance and accountable handoff. | Case and task system |
| Examples and quality tests | Show good conversation behavior and compare outcomes against unseen expert-labelled cases. | Approved prompt examples and independent test runner |
Prices, capacity and credentials stay in updateable records. The quote calculator performs arithmetic; the language model explains approved tool results. Stable conversation examples may guide the model through instructions before any model customization is considered.
Example: one question changes the quote.
Specialist’s answer: “Our monthly base covers 100 agreed accounting entries, clean records and two accounts. Each additional started block of 100 entries is extra work.”
Structured rule: a numeric allowance of 100, an increment of 100, one entity and one month, linked to BK01 and BK02.
Beth’s founder question: “About how many sales, purchase and expense entries do you handle each month? A rough count is fine; we will confirm the agreed work unit.”
| Founder’s answer | Expected behavior |
|---|---|
| 80 agreed entries, clean records, two accounts | BK01 × 1; BK02 × 0; agreed monthly period stated. |
| 280 agreed entries, clean records, two accounts | BK01 × 1; BK02 × 2; numeric calculator uses approved hours, role costs and policy. |
| Not sure | entrycount remains unknown; ask one useful clarification or show a clearly bounded, reviewable estimate. Do not infer a count from revenue. |
| Records several months behind | Assess a separately scoped cleanup phase; do not silently include unlimited historical reconstruction in a monthly fee. |
Illustrative future contract using existing pilot service IDs; not an approved Babylon rule. The current engine hardcodes some allowances, so it must be connected to structured rules before edited definitions can reliably change scope. Applicable tax or other services are separate decisions.
Review and publish a complete version.
- Captured
Keep the original answer and respondent attribution. Unknowns stay unknown.
- Structured draft
Extract typed candidates. Reject missing units, periods, source scope and invented defaults.
- Needs resolution
Conflicting definitions, currencies or rules remain unpublished; compare like segments before resolving.
- Reviewed
Service owner reviews scope, finance reviews costs, expert reviews applicability and citations.
- Approved candidate
Named owners sign off. Keep estimated and measured inputs visibly distinct.
- Tested release
Run the reference cases and critical checks. Persist the full version bundle.
- Published
Beth uses only this approved version. Keep the previous version available for rollback.
- Review due / superseded
Changed source, credential, policy or measured outcome creates a new draft; never overwrite history silently.
Who decides?
Service owners approve units, scope and deliverables. Finance approves cost and commercial policy. Experts approve applicability, citations and advice boundaries. Operations owns capacity and handoff.
What stays separate?
Keep original answers, evidence, extracted candidates and approval decisions separately. Identity and private payout data belong in controlled records. The founder conversation receives only relevant approved guidance and scoped tool results.
Proposed record types and relationships
InterviewSession
id, respondent_id, verified_role, interviewer_id, service_segment_scope, started_at, submitted_at, answer_revisions, response_status
CandidateInput
id, input_type, source_session_id, source_question_id, exact_answer_reference, typed_value, unit, period, entity_segment, evidence_ids, unknown_fields, conflict_group, valid_from, valid_until, proposed_by, review_status
Evidence
id, kind, source_reference, authority, clause, effective_date, checked_date, anonymised_job_id, estimate_or_measured, superseded_by
ReviewDecision
candidate_id, reviewer_id, verified_role, decision, reason, reviewed_at, resolved_conflicts
BethRelease
id, service_dictionary_version, package_version, rule_version, pricing_version, source_version, prompt_example_version, approved_by, published_at, previous_release_id
ReferenceScenario
id, underlying_case_group, journey, input, expected_facts, expected_clarification, expected_service_scope, reference_fee_or_bound, citation_requirements, required_escalation, prohibited_claims, expert_labels, development_or_holdout, synthetic_or_actual
TestResult
scenario_id, release_id, model_and_prompt_version, run_id, actual_facts, actual_scope, actual_quote, question_count, citation_support, critical_failures, human_review, paired_comparison
JobOutcome
anonymised_job_id, release_id, service_entity_period, planned_hours, actual_hours_by_role, realised_cost, scope_changes, quality_result, price_override_reason, calibration_review
Each approved artifact links back to the interview session, question, answer and evidence. A release pins the service, rule, pricing, source and example versions; reverting a release restores the previous complete bundle.
Audit all 24 interview topics.
Prefill known directory and catalogue details. Ask stakeholders to confirm or resolve them. Capture rules and numbers in structured fields, and obtain dynamic availability and outcome evidence from operations.
Prove the answers help before expanding.
Start with three journeys: everyday monthly records, catch-up / missing records and business setup scope.
Suggested starting panel: three delivery specialists, one finance/commercial owner and one expert. Planning numbers only; they do not establish statistical sufficiency.
Build 45 expert-labelled cases
For each journey, use typical, more complex, uncertain, conflicting-information and boundary cases. Draft three distinct cases per category. Use 30 for development and keep 15 unseen for the comparison.
Split by underlying case, keeping paraphrases together. Mark synthetic cases. Reference answers, required facts and acceptable scope are approved by humans; an AI-generated answer is not automatically the correct label.
Compare three versions
A: current Beth.
B: the same model with approved interview-derived records, rules and examples.
C: B with one input family removed at a time.
Keep the model, prompt budget, tools and source snapshot consistent. Repeat the held-out cases three times and blind the reviewer to the version.
Measure useful outcomes
- Missed necessary work
- Unnecessary service recommendations
- Correct fact interpretation and clarification
- Question burden / founder ability to answer
- Scope and fee match to the approved reference
- Citation support and applicability
- Correct specialist eligibility / escalation
- Number and severity of critical failures
Proposed release checks: no critical failures in the suite; start with a 90% expert-approved scope target and no regression in unnecessary services or question burden. Stakeholders must agree targets before running the study.
Passing a small suite is a release check, not a guarantee of production accuracy. No improvement scores are claimed in this report. Compare pricing to reviewed calculations first and actual job costs later. Sales conversion alone cannot establish correctness.
Learning after real jobs: record planned versus actual hours by role, realised cost, record quality, scope changes and delivery outcomes. The team reviews whether the rule or estimate was wrong before publishing a new version.
Model training is a later, evidence-based choice.
Use approved knowledge retrieval, business rules, quote tools and good examples first. If a persistent behavior failure remains, evaluate model customization against the same unseen cases. Dynamic legal guidance, partner costs and availability should remain maintainable records.
The OpenAI fine-tuning documentation currently says the platform is winding down and unavailable to new users. This design therefore depends on neither new fine-tuning jobs nor the hosted Evals platform; the reference cases and comparisons can run independently.
Official documentation informing the approach
- OpenAI model optimization
Supports measuring a baseline and comparing prompts/context before optional customization.
- OpenAI evaluation best practices
Supports task-specific expert-labelled evaluation and ongoing comparison. The study here is independently runnable and does not depend on the hosted Evals platform.
- OpenAI retrieval
Supports searching external knowledge and supplying relevant material to the model.
- OpenAI fine-tuning best practices
Current page says the fine-tuning platform is winding down and unavailable to new users. This design does not depend on a new fine-tuning job.