Beth’s connected stack and living knowledge
The original stack needs three clear additions: AI-quality evaluation, channel / human support, and a governed knowledge-management workspace. Keep a single owner for each responsibility. Some products share features; avoiding duplicate ownership is more realistic than expecting zero feature overlap.
Recommended: Langflow for AI steps; Temporal for durable work; LiteLLM for model routing and spend; Langfuse for AI quality; Chatwoot for messaging and human inboxes; a Beth Admin workspace backed by versioned PostgreSQL knowledge. Add operational metrics and logs through OpenTelemetry and Grafana’s stack.
Component ownership and overlap
| Component | Decision | Owns this responsibility | Boundary that prevents duplication | Official reference |
|---|---|---|---|---|
| CopilotKit + AG-UI | Keep for the website | Interactive chat and quote cards; a web UI and its event protocol. | Do not use AG-UI as the transport for WhatsApp, Telegram or LINE. | AG-UI |
| Chatwoot | Add for channel + human support | Channel inboxes, messages, agent assignment and bot-to-human transitions. | Beth remains the bot. Disable competing auto-answer bots; case fulfilment remains in Beth / Temporal. | Supported channels |
| Beth API / validated tools | Keep as application authority | Account access, consent, case records, pricing, matching, approved knowledge and channel adapters. | No framework decides legal applicability, quoted amounts or who is allowed to publish. | Existing engine scope |
| Langflow | Keep one AI flow runtime | Build profiling, research and explanation steps visually. | No parallel Dify, Flowise or LangGraph orchestration by default. Small stage flows, not a second business lifecycle. | Execution API |
| Temporal | Keep for durable business work | Recover long jobs; wait for human approval; track appointments, document requests and release stages. | A short chat turn can run through the API directly. Only Temporal owns durable Beth progression; other queues remain platform-internal. | Workflow semantics |
| LiteLLM | Promote to recommended gateway | Task-to-model routing, model access, spend, rate limits and provider fallback. | One owner of model-call retry / fallback policy. Do not duplicate it in every flow and provider SDK. | Routing controls |
| Langfuse | Add as AI quality platform | AI traces, evaluated examples, prompt versions, comparisons and feedback. | Not another agent or the authoritative tax database. Langflow owns flow code; Langfuse owns released prompt versions and evaluation records. | Evaluation loop |
| OpenTelemetry + Grafana / Prometheus / Loki | Add operational monitoring | Standard telemetry collection, system dashboards, health metrics, logs and alerting. | Langfuse owns AI-quality traces; operations dashboards own service health. Add Tempo only if operational distributed traces need their own backend. | Collector / Loki |
| PostgreSQL + object storage | Keep and make separation explicit | Versioned facts, client and service records; separate platform databases; source files and attachments in object storage. | One knowledge authority. Graph and search indexes are derived, version-checked projections; traces are not knowledge. | Relational traversal |
| Redis / tested Valkey alternative | Keep where required | Langflow events/cache, Chatwoot background jobs and observability queue/cache dependencies. | Separate durable-ish non-evicting job pools from evicting caches. Database indexes alone do not isolate memory pressure or eviction. | Langflow shared queue / Chatwoot dependencies |
| Beth Admin; Appsmith optional | Add the management experience | Interview inputs, knowledge review, service rules, source changes and publishing. | Prefer extending the existing authenticated Beth UI. Use Appsmith if faster internal UI delivery justifies another platform; it writes through the Beth API. | Appsmith |
| n8n | Defer from the core path | Optional CRM sync, email and back-office connectors. | Avoid another owner of messages, appointments, legal approval or the knowledge-release workflow. | Licence |
| Neo4j / Graphiti / pgvector | Consider after measuring need | Complex relationship traversal / time-aware graph extraction / semantic FAQ lookup, respectively. | Not three mandatory engines. Start with approved PostgreSQL relationships; do not treat model-inferred graph edges as verified tax law. | Neo4j / Graphiti |
| Docling | Use if official PDF parsing needs it | Extract structured PDF / scanned-document text for candidate evidence. | OCR and extracted tables still require source/page checks. It is a parser, not a verifier or source authority. | Docling |
How the complete path works
Website conversation
- CopilotKit + AG-UI display chat and structured cards.
- Beth API authenticates, stores the message and loads applicable published knowledge.
- Langflow runs the required AI step; LiteLLM chooses an approved model.
- Beth tools fetch evidence, calculate prices or match specialists.
- The API saves the confirmed result and streams it back to the website.
WhatsApp, Telegram or LINE
- Platform sends the message to its Chatwoot channel.
- Chatwoot’s AgentBot webhook passes it to the Beth channel adapter.
- The adapter deduplicates it and calls the same Beth API / AI / tool path.
- Beth sends the response through Chatwoot’s message API.
- Chatwoot delivers it on the original channel; human replies use that inbox.
Mobile messaging receives completed messages or platform-supported interactions, rather than the website’s token stream and AG-UI cards. Rich quote cards can become readable text plus a secure Beth portal link. Citations must remain usable across channels.
Jobs that need persistence
Long research, quote approval, appointments, document collection, escalations and knowledge releases run as Temporal workflows. Their activities call bounded Langflow stages and Beth tools. Store stage outputs and stable action IDs. Chatwoot and Langfuse have their own internal queues; those serve their platforms and do not become a second owner of Beth’s case lifecycle. A conversation being “resolved” in Chatwoot does not mean the accountant’s work is complete.
At the gateway boundary, LiteLLM owns model retry and fallback policy. Set an overall request deadline and a small provider-call retry budget; Temporal retries only an appropriate failed stage. Disable or cap overlapping flow / SDK retries to prevent a multiplication of attempts, latency and cost. LiteLLM routing / Temporal execution.
Three questions, three observability roles
| Question | Primary owner | Useful measurements |
|---|---|---|
| Which model ran, and what did it cost? | LiteLLM | Model/provider, tokens, cost, latency, failed calls, fallback and budget consumption. |
| Was the answer or decision useful and correct? | Langfuse + Beth evaluators | Evidence support, citation accuracy, applicable tax period, profile completeness, quote-rule correctness, expert corrections and handoff quality. |
| Is the service working reliably? | OpenTelemetry + Prometheus / Loki / Grafana | API latency, webhook failures, worker health, backlog, cache freshness, source errors, failed notifications, storage and recovery alerts. |
LiteLLM has logging integrations, and Langflow has a documented Langfuse integration. That helps, but a gateway-only trace misses profiling, source fetches, pricing, approvals and message delivery. Instrument the Beth API and tools, propagate the parent trace through workers and HTTP calls, and preserve a shared request / conversation / case correlation ID. LiteLLM logging / Langflow integration.
Choose one export path per span. Avoid sending the same LLM call to Langfuse through both a gateway callback and a second automatic flow instrumenter. Preserve the root span and parent context; duplicate observations can distort metrics. The operational collector processes / exports telemetry; it does not itself provide durable storage or dashboards. Langfuse ingestion guidance / Collector role.
Langfuse is my preferred AI-quality platform here. Its open-source core covers the evaluation workflow Beth needs. LangSmith is a credible alternative, but operating both would duplicate much of the tracing, datasets and prompt-management work. Self-hosted LangSmith is an Enterprise add-on rather than an open-source platform requirement. Langfuse evaluations / LangSmith self-hosting.
Telemetry is not the permanent client record or legal audit trail. Store durable approval and case events in Beth’s database, redact client content before sending traces, and use retention policies. Do not put raw personal identifiers into high-cardinality metric labels. Audit monetary totals against provider billing; a reported token estimate is not an invoice.
Channel delivery and live specialist attendance
Chatwoot documents WhatsApp, Telegram and LINE channels, plus an API channel suitable for custom applications. Use Beth as an external AgentBot through webhooks and REST APIs; do not add Chatwoot Captain as a second answering brain. Captain also has different edition / commercial boundaries. Channels / AgentBot integration / Self-hosted editions.
| Channel | Setup / delivery requirements | Beth integration responsibility |
|---|---|---|
| WhatsApp Business account, approved number, Cloud API configuration and templates when required outside the customer-service window. Platform charges may apply. | Use Chatwoot’s channel connector; handle receipts, consent, template eligibility and failed sends. | |
| Telegram | A dedicated BotFather-created bot and channel credentials; normal Bot API and Business mode have different behaviour. | Deduplicate provider updates; verify authenticated webhook delivery. Do not assume every specialist online on Telegram can automatically take over Beth. |
| LINE | A LINE Official Account and Messaging API credentials; verify signatures and handle provider redelivery / reply-token constraints. | Use Chatwoot’s connector, prove attachment and outbound behaviour for the selected version; do not wait on slow AI work before acknowledging the webhook. |
| Beth website | Keep the existing branded CopilotKit experience. | Bridge human conversations to a Chatwoot API inbox, keeping one transcript identity rather than adding a competing website widget. |
Channel references: LINE signature verification / LINE receiving / redelivery / Telegram API / Chatwoot Telegram setup.
The specialist’s primary shared inbox should be Chatwoot; the jobs-in-progress portal remains Beth’s case-management surface. Asking for a human opens the conversation for a specialist and records a Beth ticket. Suspend automatic bot replies while a human owns the conversation, including any AI response already in flight. Offline specialists receive notifications and pick up the saved case later.
If specialists must reply from their personal Telegram interface, add an explicit authenticated staff relay: verify staff identity and permissions, select the case, relay only to the correct founder and record the event. Presence is separate from permission and capacity. Channel identities are linked to a Beth account only after a verified linking step; never merge people merely because their display names match.
Give staff a Beth Admin workspace
The central management tool is a business-facing Beth workspace. Specialists answer guided interview questions; editors submit source changes; reviewers approve evidence; managers approve pricing. Langflow is for technical flow authors, and Langfuse is for evaluation / prompt work. Neither is the main legislation and service-rule management screen.
Prefer extending the authenticated Beth admin / interview UI already being built. If rapid low-code internal delivery is the priority, Appsmith is a useful alternative UI builder. Its core is Apache 2.0 and it can call REST APIs. It is not the knowledge authority. Granular Appsmith access control is documented as a Business feature, so Community Edition alone must not be assumed to provide every reviewer / publisher permission. Appsmith core / API connections / Access-control boundary.
Every action goes through the Beth API. Authenticate the real end user server-side and enforce author / reviewer / publisher roles there. A global admin API key, a hidden button or a client-supplied email is not an adequate approval identity. Prove the selected UI’s per-user authentication bridge before choosing it; if that cannot be done cleanly, use the existing Beth UI or a suitable commercial identity integration.
| Admin area | Information collected | How Beth uses it |
|---|---|---|
| Interview and service editor | Service code, problem it solves, qualifying answers, included work, exclusions, effort range, volume drivers, frequency, prerequisites and exceptions. | Profile rules, package composition, reproducible quote calculations and specialist requirements. |
| Legislation / source editor | Official URL, document identity, code section / page, exact supporting excerpt, issuing authority, publication / effective dates, applicable period and amendment links. | Evidence retrieval, citations, applicability filters and impact checks. |
| Change comparison | Original vs proposed fact / rule; who proposed it; affected FAQ, services, questions and test cases. | Reviewer can see consequences rather than approve an isolated upload. |
| Review and release desk | Expert decision, regression results, manager approval, release version, effective date and rollback reason. | Publish approved changes, retire superseded facts and record an audit trail. |
| Quality inbox | Flagged answers, redacted conversation samples, expert corrections, source failures and quote overrides. | Create reviewed examples and test improvements before releasing them. |
The existing questionnaire should save structured data, not just long free-text answers. Each question must map to a field or rule and at least one acceptance example. An answer that changes neither a decision nor an evaluation case is a candidate for removal from the interview.
Where the living brain sits
The Knowledge Service sits behind Beth’s tools and between Langflow’s questions and the source evidence / business rules it is allowed to use. PostgreSQL stores approved facts, relationships, applicability and versions. Bounded object storage preserves necessary source excerpts / files; an optional graph or search index is a derived view, not a second truth store. The service remains small and focused; Beth still fetches current official material when a question requires it.
Inputs become reviewed knowledge, then feedback becomes reviewed tests.
Open the knowledge-loop diagramFour stores with different purposes
| Store | Purpose | Update boundary |
|---|---|---|
| Shared approved knowledge | Legal propositions, FAQ, service relationships and source evidence. | Expert-reviewed; versioned; applicable dates and authority recorded. |
| Business configuration | Question rules, service codes, effort models, pricing and specialist skills. | Operational / manager approval; deterministic validation; existing accepted quotes are not silently repriced. |
| Private founder memory | A founder’s own profile, conversations, documents and cases. | Account-isolated; corrected by authorised users; never promoted into shared legal knowledge automatically. |
| Evaluation memory | Reviewed examples, expected answers and outcomes. | Redacted and reviewed; used for tests and analysis, not directly retrieved as legal evidence. |
Relationships make the brain useful
Start with typed links: source supports fact; new provision amends / supersedes old provision; fact applies to business category and tax period; obligation requires service item; package includes service item; service needs specialist skill; question determines a qualifying condition. Those links allow Beth to explain a recommendation and let staff locate everything affected by a change.
Each fact and relationship needs: a stable ID; statement / typed condition; evidence source and excerpt; business / geographic scope; effective-from and effective-to dates; recorded-at and retirement timestamps; publication and last-verification dates; approval state and reviewer; version; and superseded-by / conflict links. Model confidence alone cannot establish truth. Keep relationships grounded in approved evidence.
The two timelines matter: when a rule applies and when Beth learned / approved it. A rule retired from current advice may still apply to a founder’s earlier filing period. A future-effective rule may be published now without applying today. A changed page or new circular does not automatically repeal every older provision.
See current versus historical retrieval
Additions and removals
- Capture: preserve the source, its checksum, dates and relevant excerpt. Parse it into proposed facts / links. Unverified extraction remains draft.
- Compare: identify additions, changes, contradictions and affected questions / packages. A source change triggers revalidation rather than proving that the legal meaning changed.
- Review and test: an expert checks authority, scope, dates and interpretation; test expected outcomes. Quarantine unresolved conflicts and suppress affected claims when necessary.
- Publish: atomically activate an immutable release manifest with exact fact, rule, prompt and flow versions. Preserve old versions and approval events.
- Invalidate: retire superseded items from current retrieval; invalidate dependent cache keys, embeddings and graph projections. Queries verify the active version so a delayed projection cannot leak old advice.
- Retain or purge appropriately: preserve legal / decision history where needed. Private client data follows its own access, retention and deletion policy; it is not kept forever merely because facts are versioned.
For requests that require fresh law, fetch authoritative sources at answer time and show the checked-at date. If those sites are unavailable, use approved alternate access or clearly dated verified material; do not silently claim a stale answer is current. A bounded cache does not guarantee current law.
Do we need a dedicated graph database?
Not at the beginning. PostgreSQL tables for facts and typed edges, plus recursive queries, can implement the initial curated graph. Add Neo4j when measured query / editing needs justify a dedicated graph engine. Neo4j Community is GPLv3; enterprise capabilities have separate commercial terms. A graph adds retrieval power, not legal validity. PostgreSQL recursive queries / Neo4j editions.
Graphiti is an open-source candidate for time-aware relationship extraction and memory. Use it only in the candidate / derived-memory path initially: extraction and contradiction handling do not replace a tax expert or the approved registry. Do not automatically feed private chats into a shared graph. Graphiti’s “temporal” knowledge refers to time-aware facts; it does not replace Temporal’s durable workflows. Graphiti source and design.
How Beth gets smarter safely
- Langfuse records redacted traces, feedback and quality scores.
- Experts review failures and add expected answers / outcomes to a dataset.
- Editors propose prompt, flow, knowledge or pricing-rule changes.
- Tests compare the proposed release with the current version.
- An authorised reviewer publishes the release; monitoring checks its effect.
Use deterministic checks for quote arithmetic, source IDs, required fields and dates. Use experts for tax interpretation. An LLM judge can triage answer quality, but calibrate it against expert labels and do not give it final legal approval. Keep a held-out test set to avoid calling memorised examples an improvement. Langfuse evaluation methods.
Real-time updating is achievable; real-time weight training is a different mechanism. Staff can edit and publish knowledge / rules without redeploying the whole app. Next requests read the published version. Active sessions receive a knowledge-change notice when relevant; in-flight stages must revalidate materially affected advice rather than silently changing their reasoning. Accepted quotes keep their priced scope and version unless explicitly renegotiated.
Define and measure a publication-to-visibility target. “Instant” is not automatic: Langfuse prompt SDKs cache prompts, and indexes / worker caches can lag. Resolve exact prompt versions from the release manifest, invalidate caches and fall back to the registry when projections are behind. Flow / workflow code changes still go through deployment and compatibility checks. Prompt cache behaviour.
The overhead hidden behind engine names
- Langfuse self-hosting: web + worker, PostgreSQL, ClickHouse, Redis / Valkey and S3-compatible object storage. It is not just another table in Beth’s database. Managed hosting is an option after a data-handling review. Deployment architecture.
- Chatwoot: application services, Sidekiq workers, PostgreSQL, Redis, attachments and channel credentials. Its internal message-delivery jobs remain its responsibility. System requirements.
- Appsmith, if selected: another admin-platform deployment with its own state, including MongoDB and Redis in the documented deployment architecture. Extending the existing Beth UI avoids that extra platform. Appsmith architecture.
- Temporal: server services, workers, persistence, deployment versioning and recovery operations. Run version-compatible workers when changing long-lived workflows.
- Identity and storage: use the existing Babylon identity if it meets requirements; consider Keycloak only if SSO / MFA needs justify it. Separate databases, roles, storage buckets, permissions, backups and restore exercises. Shared infrastructure capacity does not mean shared credentials.
- Open-source editions: verify required features and licences for pinned versions. Langfuse and Chatwoot have commercial features; LangSmith self-hosting is commercial; n8n is source-available. Redis terms vary by version. Hosted AI and messaging accounts are external services with their own costs.
Do not add a second full RAG platform, standalone vector database, graph database, agent framework or workflow engine just because it offers a feature already assigned elsewhere. Search relevance, graph value and hardware requirements should be measured with Beth’s actual workload before expansion.
What proves the end-to-end design?
These are acceptance scenarios for the future integration, not tests already passed by the local demo.
| Proof scenario | Expected evidence |
|---|---|
| One question on each channel | Same Beth API and approved versions; source citations delivered; no cross-account transcript leakage. |
| Webhook arrives twice / out of order | One logical message / case action; authenticated provider receipt; ordering policy recorded. |
| Human takes over while AI is generating | Only the human’s permitted reply is delivered; late bot result is suppressed; case remains trackable. |
| Worker crashes after an appointment succeeds | Workflow resumes; no second appointment or notification caused by retries. |
| Model fails or budget is exceeded | Approved fallback or clear failure; bounded attempts; cost accounted once per actual provider call. |
| New provision changes a relevant condition | Expert-reviewed change identifies impacted FAQ / rules; current retrieval excludes the superseded item; earlier-period query still finds applicable history. |
| Cache / graph update is delayed | Active-version check detects the stale projection and uses approved registry data or pauses affected advice. |
| Bad release or unauthorised publisher | Permission denied or rollout blocked; a rollback switches versions; recorded decisions remain attributable. |
| One founder reports a wrong answer | Trace links question, model call, tool / source and rule versions; reviewed correction becomes a test, not an automatic live rule. |
| Source site or observability platform is down | Source freshness policy governs the answer; telemetry outage alerts but does not normally block a valid response; essential audits persist in Beth. |
Implementation order: prove the Beth knowledge API and quote rules first; add LiteLLM + Langfuse instrumentation; implement one channel and human takeover; exercise Temporal recovery; then expand to additional channels and a richer admin UI. A graph engine comes after demonstrated value. Re-estimate delivery effort once deployment choices and integration proof are settled.