Babylon 2k
Research and proposed design · 6 October 2026

Beth’s connected stack and living knowledge

The original stack needs three clear additions: AI-quality evaluation, channel / human support, and a governed knowledge-management workspace. Keep a single owner for each responsibility. Some products share features; avoiding duplicate ownership is more realistic than expecting zero feature overlap.

Recommended: Langflow for AI steps; Temporal for durable work; LiteLLM for model routing and spend; Langfuse for AI quality; Chatwoot for messaging and human inboxes; a Beth Admin workspace backed by versioned PostgreSQL knowledge. Add operational metrics and logs through OpenTelemetry and Grafana’s stack.

This is a documented proposal, not an installed stack or a verified production integration. Official documentation was reviewed; exact version, edition, identity integration and recovery compatibility still need a proof of concept. The current lean scope supersedes this full-stack reference for rollout and effort. Use the revised scope for the launch sequence and decision-maker commercial review.

Component ownership and overlap

ComponentDecisionOwns this responsibilityBoundary that prevents duplicationOfficial reference
CopilotKit + AG-UIKeep for the websiteInteractive chat and quote cards; a web UI and its event protocol.Do not use AG-UI as the transport for WhatsApp, Telegram or LINE.AG-UI
ChatwootAdd for channel + human supportChannel inboxes, messages, agent assignment and bot-to-human transitions.Beth remains the bot. Disable competing auto-answer bots; case fulfilment remains in Beth / Temporal.Supported channels
Beth API / validated toolsKeep as application authorityAccount access, consent, case records, pricing, matching, approved knowledge and channel adapters.No framework decides legal applicability, quoted amounts or who is allowed to publish.Existing engine scope
LangflowKeep one AI flow runtimeBuild profiling, research and explanation steps visually.No parallel Dify, Flowise or LangGraph orchestration by default. Small stage flows, not a second business lifecycle.Execution API
TemporalKeep for durable business workRecover long jobs; wait for human approval; track appointments, document requests and release stages.A short chat turn can run through the API directly. Only Temporal owns durable Beth progression; other queues remain platform-internal.Workflow semantics
LiteLLMPromote to recommended gatewayTask-to-model routing, model access, spend, rate limits and provider fallback.One owner of model-call retry / fallback policy. Do not duplicate it in every flow and provider SDK.Routing controls
LangfuseAdd as AI quality platformAI traces, evaluated examples, prompt versions, comparisons and feedback.Not another agent or the authoritative tax database. Langflow owns flow code; Langfuse owns released prompt versions and evaluation records.Evaluation loop
OpenTelemetry + Grafana / Prometheus / LokiAdd operational monitoringStandard telemetry collection, system dashboards, health metrics, logs and alerting.Langfuse owns AI-quality traces; operations dashboards own service health. Add Tempo only if operational distributed traces need their own backend.Collector / Loki
PostgreSQL + object storageKeep and make separation explicitVersioned facts, client and service records; separate platform databases; source files and attachments in object storage.One knowledge authority. Graph and search indexes are derived, version-checked projections; traces are not knowledge.Relational traversal
Redis / tested Valkey alternativeKeep where requiredLangflow events/cache, Chatwoot background jobs and observability queue/cache dependencies.Separate durable-ish non-evicting job pools from evicting caches. Database indexes alone do not isolate memory pressure or eviction.Langflow shared queue / Chatwoot dependencies
Beth Admin; Appsmith optionalAdd the management experienceInterview inputs, knowledge review, service rules, source changes and publishing.Prefer extending the existing authenticated Beth UI. Use Appsmith if faster internal UI delivery justifies another platform; it writes through the Beth API.Appsmith
n8nDefer from the core pathOptional CRM sync, email and back-office connectors.Avoid another owner of messages, appointments, legal approval or the knowledge-release workflow.Licence
Neo4j / Graphiti / pgvectorConsider after measuring needComplex relationship traversal / time-aware graph extraction / semantic FAQ lookup, respectively.Not three mandatory engines. Start with approved PostgreSQL relationships; do not treat model-inferred graph edges as verified tax law.Neo4j / Graphiti
DoclingUse if official PDF parsing needs itExtract structured PDF / scanned-document text for candidate evidence.OCR and extracted tables still require source/page checks. It is a parser, not a verifier or source authority.Docling

How the complete path works

Website conversation

  1. CopilotKit + AG-UI display chat and structured cards.
  2. Beth API authenticates, stores the message and loads applicable published knowledge.
  3. Langflow runs the required AI step; LiteLLM chooses an approved model.
  4. Beth tools fetch evidence, calculate prices or match specialists.
  5. The API saves the confirmed result and streams it back to the website.

WhatsApp, Telegram or LINE

  1. Platform sends the message to its Chatwoot channel.
  2. Chatwoot’s AgentBot webhook passes it to the Beth channel adapter.
  3. The adapter deduplicates it and calls the same Beth API / AI / tool path.
  4. Beth sends the response through Chatwoot’s message API.
  5. Chatwoot delivers it on the original channel; human replies use that inbox.

Mobile messaging receives completed messages or platform-supported interactions, rather than the website’s token stream and AG-UI cards. Rich quote cards can become readable text plus a secure Beth portal link. Citations must remain usable across channels.

Jobs that need persistence

Long research, quote approval, appointments, document collection, escalations and knowledge releases run as Temporal workflows. Their activities call bounded Langflow stages and Beth tools. Store stage outputs and stable action IDs. Chatwoot and Langfuse have their own internal queues; those serve their platforms and do not become a second owner of Beth’s case lifecycle. A conversation being “resolved” in Chatwoot does not mean the accountant’s work is complete.

At the gateway boundary, LiteLLM owns model retry and fallback policy. Set an overall request deadline and a small provider-call retry budget; Temporal retries only an appropriate failed stage. Disable or cap overlapping flow / SDK retries to prevent a multiplication of attempts, latency and cost. LiteLLM routing / Temporal execution.

Three questions, three observability roles

QuestionPrimary ownerUseful measurements
Which model ran, and what did it cost?LiteLLMModel/provider, tokens, cost, latency, failed calls, fallback and budget consumption.
Was the answer or decision useful and correct?Langfuse + Beth evaluatorsEvidence support, citation accuracy, applicable tax period, profile completeness, quote-rule correctness, expert corrections and handoff quality.
Is the service working reliably?OpenTelemetry + Prometheus / Loki / GrafanaAPI latency, webhook failures, worker health, backlog, cache freshness, source errors, failed notifications, storage and recovery alerts.

LiteLLM has logging integrations, and Langflow has a documented Langfuse integration. That helps, but a gateway-only trace misses profiling, source fetches, pricing, approvals and message delivery. Instrument the Beth API and tools, propagate the parent trace through workers and HTTP calls, and preserve a shared request / conversation / case correlation ID. LiteLLM logging / Langflow integration.

Choose one export path per span. Avoid sending the same LLM call to Langfuse through both a gateway callback and a second automatic flow instrumenter. Preserve the root span and parent context; duplicate observations can distort metrics. The operational collector processes / exports telemetry; it does not itself provide durable storage or dashboards. Langfuse ingestion guidance / Collector role.

Langfuse is my preferred AI-quality platform here. Its open-source core covers the evaluation workflow Beth needs. LangSmith is a credible alternative, but operating both would duplicate much of the tracing, datasets and prompt-management work. Self-hosted LangSmith is an Enterprise add-on rather than an open-source platform requirement. Langfuse evaluations / LangSmith self-hosting.

Telemetry is not the permanent client record or legal audit trail. Store durable approval and case events in Beth’s database, redact client content before sending traces, and use retention policies. Do not put raw personal identifiers into high-cardinality metric labels. Audit monetary totals against provider billing; a reported token estimate is not an invoice.

Channel delivery and live specialist attendance

Chatwoot documents WhatsApp, Telegram and LINE channels, plus an API channel suitable for custom applications. Use Beth as an external AgentBot through webhooks and REST APIs; do not add Chatwoot Captain as a second answering brain. Captain also has different edition / commercial boundaries. Channels / AgentBot integration / Self-hosted editions.

ChannelSetup / delivery requirementsBeth integration responsibility
WhatsAppWhatsApp Business account, approved number, Cloud API configuration and templates when required outside the customer-service window. Platform charges may apply.Use Chatwoot’s channel connector; handle receipts, consent, template eligibility and failed sends.
TelegramA dedicated BotFather-created bot and channel credentials; normal Bot API and Business mode have different behaviour.Deduplicate provider updates; verify authenticated webhook delivery. Do not assume every specialist online on Telegram can automatically take over Beth.
LINEA LINE Official Account and Messaging API credentials; verify signatures and handle provider redelivery / reply-token constraints.Use Chatwoot’s connector, prove attachment and outbound behaviour for the selected version; do not wait on slow AI work before acknowledging the webhook.
Beth websiteKeep the existing branded CopilotKit experience.Bridge human conversations to a Chatwoot API inbox, keeping one transcript identity rather than adding a competing website widget.

Channel references: LINE signature verification / LINE receiving / redelivery / Telegram API / Chatwoot Telegram setup.

The specialist’s primary shared inbox should be Chatwoot; the jobs-in-progress portal remains Beth’s case-management surface. Asking for a human opens the conversation for a specialist and records a Beth ticket. Suspend automatic bot replies while a human owns the conversation, including any AI response already in flight. Offline specialists receive notifications and pick up the saved case later.

If specialists must reply from their personal Telegram interface, add an explicit authenticated staff relay: verify staff identity and permissions, select the case, relay only to the correct founder and record the event. Presence is separate from permission and capacity. Channel identities are linked to a Beth account only after a verified linking step; never merge people merely because their display names match.

Give staff a Beth Admin workspace

The central management tool is a business-facing Beth workspace. Specialists answer guided interview questions; editors submit source changes; reviewers approve evidence; managers approve pricing. Langflow is for technical flow authors, and Langfuse is for evaluation / prompt work. Neither is the main legislation and service-rule management screen.

Prefer extending the authenticated Beth admin / interview UI already being built. If rapid low-code internal delivery is the priority, Appsmith is a useful alternative UI builder. Its core is Apache 2.0 and it can call REST APIs. It is not the knowledge authority. Granular Appsmith access control is documented as a Business feature, so Community Edition alone must not be assumed to provide every reviewer / publisher permission. Appsmith core / API connections / Access-control boundary.

Every action goes through the Beth API. Authenticate the real end user server-side and enforce author / reviewer / publisher roles there. A global admin API key, a hidden button or a client-supplied email is not an adequate approval identity. Prove the selected UI’s per-user authentication bridge before choosing it; if that cannot be done cleanly, use the existing Beth UI or a suitable commercial identity integration.

Admin areaInformation collectedHow Beth uses it
Interview and service editorService code, problem it solves, qualifying answers, included work, exclusions, effort range, volume drivers, frequency, prerequisites and exceptions.Profile rules, package composition, reproducible quote calculations and specialist requirements.
Legislation / source editorOfficial URL, document identity, code section / page, exact supporting excerpt, issuing authority, publication / effective dates, applicable period and amendment links.Evidence retrieval, citations, applicability filters and impact checks.
Change comparisonOriginal vs proposed fact / rule; who proposed it; affected FAQ, services, questions and test cases.Reviewer can see consequences rather than approve an isolated upload.
Review and release deskExpert decision, regression results, manager approval, release version, effective date and rollback reason.Publish approved changes, retire superseded facts and record an audit trail.
Quality inboxFlagged answers, redacted conversation samples, expert corrections, source failures and quote overrides.Create reviewed examples and test improvements before releasing them.

The existing questionnaire should save structured data, not just long free-text answers. Each question must map to a field or rule and at least one acceptance example. An answer that changes neither a decision nor an evaluation case is a candidate for removal from the interview.

Where the living brain sits

The Knowledge Service sits behind Beth’s tools and between Langflow’s questions and the source evidence / business rules it is allowed to use. PostgreSQL stores approved facts, relationships, applicability and versions. Bounded object storage preserves necessary source excerpts / files; an optional graph or search index is a derived view, not a second truth store. The service remains small and focused; Beth still fetches current official material when a question requires it.

Inputs become reviewed knowledge, then feedback becomes reviewed tests.

Open the knowledge-loop diagram

Four stores with different purposes

StorePurposeUpdate boundary
Shared approved knowledgeLegal propositions, FAQ, service relationships and source evidence.Expert-reviewed; versioned; applicable dates and authority recorded.
Business configurationQuestion rules, service codes, effort models, pricing and specialist skills.Operational / manager approval; deterministic validation; existing accepted quotes are not silently repriced.
Private founder memoryA founder’s own profile, conversations, documents and cases.Account-isolated; corrected by authorised users; never promoted into shared legal knowledge automatically.
Evaluation memoryReviewed examples, expected answers and outcomes.Redacted and reviewed; used for tests and analysis, not directly retrieved as legal evidence.

Relationships make the brain useful

Start with typed links: source supports fact; new provision amends / supersedes old provision; fact applies to business category and tax period; obligation requires service item; package includes service item; service needs specialist skill; question determines a qualifying condition. Those links allow Beth to explain a recommendation and let staff locate everything affected by a change.

Each fact and relationship needs: a stable ID; statement / typed condition; evidence source and excerpt; business / geographic scope; effective-from and effective-to dates; recorded-at and retirement timestamps; publication and last-verification dates; approval state and reviewer; version; and superseded-by / conflict links. Model confidence alone cannot establish truth. Keep relationships grounded in approved evidence.

The two timelines matter: when a rule applies and when Beth learned / approved it. A rule retired from current advice may still apply to a founder’s earlier filing period. A future-effective rule may be published now without applying today. A changed page or new circular does not automatically repeal every older provision.

See current versus historical retrieval

Illustrative records only. Rule A and Rule B are placeholders, not Philippine tax provisions.

Additions and removals

  1. Capture: preserve the source, its checksum, dates and relevant excerpt. Parse it into proposed facts / links. Unverified extraction remains draft.
  2. Compare: identify additions, changes, contradictions and affected questions / packages. A source change triggers revalidation rather than proving that the legal meaning changed.
  3. Review and test: an expert checks authority, scope, dates and interpretation; test expected outcomes. Quarantine unresolved conflicts and suppress affected claims when necessary.
  4. Publish: atomically activate an immutable release manifest with exact fact, rule, prompt and flow versions. Preserve old versions and approval events.
  5. Invalidate: retire superseded items from current retrieval; invalidate dependent cache keys, embeddings and graph projections. Queries verify the active version so a delayed projection cannot leak old advice.
  6. Retain or purge appropriately: preserve legal / decision history where needed. Private client data follows its own access, retention and deletion policy; it is not kept forever merely because facts are versioned.

For requests that require fresh law, fetch authoritative sources at answer time and show the checked-at date. If those sites are unavailable, use approved alternate access or clearly dated verified material; do not silently claim a stale answer is current. A bounded cache does not guarantee current law.

Do we need a dedicated graph database?

Not at the beginning. PostgreSQL tables for facts and typed edges, plus recursive queries, can implement the initial curated graph. Add Neo4j when measured query / editing needs justify a dedicated graph engine. Neo4j Community is GPLv3; enterprise capabilities have separate commercial terms. A graph adds retrieval power, not legal validity. PostgreSQL recursive queries / Neo4j editions.

Graphiti is an open-source candidate for time-aware relationship extraction and memory. Use it only in the candidate / derived-memory path initially: extraction and contradiction handling do not replace a tax expert or the approved registry. Do not automatically feed private chats into a shared graph. Graphiti’s “temporal” knowledge refers to time-aware facts; it does not replace Temporal’s durable workflows. Graphiti source and design.

How Beth gets smarter safely

  1. Langfuse records redacted traces, feedback and quality scores.
  2. Experts review failures and add expected answers / outcomes to a dataset.
  3. Editors propose prompt, flow, knowledge or pricing-rule changes.
  4. Tests compare the proposed release with the current version.
  5. An authorised reviewer publishes the release; monitoring checks its effect.

Use deterministic checks for quote arithmetic, source IDs, required fields and dates. Use experts for tax interpretation. An LLM judge can triage answer quality, but calibrate it against expert labels and do not give it final legal approval. Keep a held-out test set to avoid calling memorised examples an improvement. Langfuse evaluation methods.

Real-time updating is achievable; real-time weight training is a different mechanism. Staff can edit and publish knowledge / rules without redeploying the whole app. Next requests read the published version. Active sessions receive a knowledge-change notice when relevant; in-flight stages must revalidate materially affected advice rather than silently changing their reasoning. Accepted quotes keep their priced scope and version unless explicitly renegotiated.

Define and measure a publication-to-visibility target. “Instant” is not automatic: Langfuse prompt SDKs cache prompts, and indexes / worker caches can lag. Resolve exact prompt versions from the release manifest, invalidate caches and fall back to the registry when projections are behind. Flow / workflow code changes still go through deployment and compatibility checks. Prompt cache behaviour.

The overhead hidden behind engine names

Do not add a second full RAG platform, standalone vector database, graph database, agent framework or workflow engine just because it offers a feature already assigned elsewhere. Search relevance, graph value and hardware requirements should be measured with Beth’s actual workload before expansion.

What proves the end-to-end design?

These are acceptance scenarios for the future integration, not tests already passed by the local demo.

Proof scenarioExpected evidence
One question on each channelSame Beth API and approved versions; source citations delivered; no cross-account transcript leakage.
Webhook arrives twice / out of orderOne logical message / case action; authenticated provider receipt; ordering policy recorded.
Human takes over while AI is generatingOnly the human’s permitted reply is delivered; late bot result is suppressed; case remains trackable.
Worker crashes after an appointment succeedsWorkflow resumes; no second appointment or notification caused by retries.
Model fails or budget is exceededApproved fallback or clear failure; bounded attempts; cost accounted once per actual provider call.
New provision changes a relevant conditionExpert-reviewed change identifies impacted FAQ / rules; current retrieval excludes the superseded item; earlier-period query still finds applicable history.
Cache / graph update is delayedActive-version check detects the stale projection and uses approved registry data or pauses affected advice.
Bad release or unauthorised publisherPermission denied or rollout blocked; a rollback switches versions; recorded decisions remain attributable.
One founder reports a wrong answerTrace links question, model call, tool / source and rule versions; reviewed correction becomes a test, not an automatic live rule.
Source site or observability platform is downSource freshness policy governs the answer; telemetry outage alerts but does not normally block a valid response; essential audits persist in Beth.

Implementation order: prove the Beth knowledge API and quote rules first; add LiteLLM + Langfuse instrumentation; implement one channel and human takeover; exercise Temporal recovery; then expand to additional channels and a richer admin UI. A graph engine comes after demonstrated value. Re-estimate delivery effort once deployment choices and integration proof are settled.