| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
Adds six optional server env vars: - AI_BASE_URL / AI_API_KEY / AI_MODEL — chat completions provider - STT_BASE_URL / STT_API_KEY / STT_MODEL — speech-to-text provider All are .optional() so existing self-hosters running Groq/OpenAI/Deepgram are unaffected. Subsequent commits wire them into the AI client and the transcription workflow.
Replaces the groq-sdk wrapper and the hand-rolled fetch-to-OpenAI fallback with a single `openai` SDK client in apps/web/lib/ai-provider.ts. The new getAiClient()/getAiModel() resolve, in order: 1. AI_BASE_URL + AI_API_KEY + AI_MODEL (any OpenAI-compatible provider: Ollama, vLLM, OpenRouter, LiteLLM, etc.) 2. GROQ_API_KEY (existing default, baseURL pinned to Groq, model preserved as openai/gpt-oss-120b) 3. OPENAI_API_KEY (default OpenAI endpoint, model preserved as gpt-4o-mini) Call sites migrated: - apps/web/workflows/generate-ai.ts: drops the duplicate callOpenAi raw fetch in favor of the unified client; signatures threaded - apps/web/actions/videos/translate-transcript.ts - apps/web/lib/messenger/agent.ts (Groq branch only; Anthropic/OpenAI fallbacks untouched — separate domain, kept out of scope) Dependency change: -groq-sdk, +openai. Behavior for existing installs is unchanged because the Groq path now constructs an OpenAI client with baseURL = https://api.groq.com/openai/v1 — same wire protocol.
Transcription workflow: - When STT_BASE_URL is set, transcribeAudio dispatches to a new transcribeViaSttProvider() that calls openai.audio.transcriptions.create with response_format: "vtt". The OpenAI Whisper API returns WebVTT directly, which is exactly what Cap's pipeline writes to S3, so the Deepgram-specific formatToWebVTT(DeepgramResult) adapter drops out on this path. - Default behavior (STT_BASE_URL unset) still uses Deepgram. Existing installs are unaffected. Trigger-gate widenings (these were the blockers preventing self-hosters on local providers from ever firing the workflow): - apps/web/lib/transcribe.ts: accept STT_BASE_URL as a valid provider - apps/web/lib/generate-ai.ts: accept AI_BASE_URL as a valid provider - apps/web/actions/videos/get-status.ts: same widenings in the share-page auto-trigger paths for both transcription and AI generation
The default docker-compose.yml did not pass DEEPGRAM/GROQ/OPENAI env vars through to cap-web, which is part of why self-host AI was silently broken — even when users set the keys in .env, they never reached the container. This commit threads them through along with the new AI_*/STT_* triples in all four compose flavors: - docker-compose.yml (default) - docker-compose.template.yml - docker-compose.coolify.yml - docker-compose.coolify.env.example
…scheme
- AI_BASE_URL and STT_BASE_URL now refuse non-http(s) schemes (defense
against typos like `file://` or `gopher://`). Empty string still passes
for compose default `${VAR:-}`.
- Doc strings updated to make the requirement contract explicit
(AI_API_KEY and AI_MODEL are required when AI_BASE_URL is set; same for
STT_*). The previous wording said AI_API_KEY "falls back to GROQ_API_KEY
or OPENAI_API_KEY" — that fallback is removed in the next commit
because it could silently send a paid cloud key to an arbitrary URL.
…e gates
ai-provider.ts:
- Drop the module-level singleton cache. The OpenAI SDK is cheap to
construct and the cache made env changes / hot-reloads / tests carry
stale clients with no path to recreate.
- Drop the cross-provider apiKey fallback. Previously, setting AI_BASE_URL
without AI_API_KEY would silently send the configured GROQ_API_KEY or
OPENAI_API_KEY over the wire to the new endpoint. Now AI_API_KEY is
required explicitly when AI_BASE_URL is set; same for STT_API_KEY.
- Throw clear errors when AI_BASE_URL is set without the required
AI_API_KEY or AI_MODEL (and STT analogue). The previous code would
silently default AI_MODEL to "gpt-4o-mini" and let Ollama/vLLM return
an opaque 404 inside the workflow step.
- Set explicit timeouts (120s chat, 300s STT) and maxRetries: 2 on the
OpenAI client. The SDK default of 600s would hang workflow steps for
10 minutes on a stuck local inference call; the retry restores the
resilience that the previous Groq->OpenAI fallback used to provide.
- Add isAiConfigured() / isSttConfigured() helpers as the single source
of truth for "is any chat / STT provider available?" so the OR-chains
in trigger gates don't drift the next time a provider type lands.
workflows/transcribe.ts:
- Drop the `as unknown as string` cast on the OpenAI SDK transcription
response. With `response_format: "vtt" as const` the SDK's overload
narrows to string at compile time; the unsafe cast was hiding that.
- Strengthen the WebVTT smoke check from a substring search for
"WEBVTT" to a structural check (`/^WEBVTT/m` header line plus a cue
arrow `-->`). The substring form would both reject valid VTT without
the header and accept SRT or other formats that happened to contain
the word.
workflows/generate-ai.ts:
- Request `response_format: { type: "json_object" }` on chat completions.
Every prompt already instructs "Return ONLY valid JSON"; modern
OpenAI-compatible providers (OpenAI, Groq, Ollama, vLLM, OpenRouter)
enforce that with this flag, which materially reduces parse failures
on smaller local models. A try/catch falls back to plain mode when
the underlying provider rejects the field, keeping niche gateways
compatible.
Trigger gates consolidated via isAiConfigured() / isSttConfigured():
- apps/web/actions/videos/get-status.ts (share-page auto-trigger, both
transcription and AI-generation paths)
- apps/web/lib/transcribe.ts (lib entry point)
- apps/web/lib/generate-ai.ts (lib entry point)
- apps/web/workflows/generate-ai.ts (validateAndSetProcessing step;
also aligns the second-check error message with the first)
- apps/web/workflows/transcribe.ts (validateVideo step)
The previous PR commit dropped the cross-provider failover and described maxRetries: 2 as a replacement — that was wrong. maxRetries only retries the same endpoint on transient errors; it does not preserve the prior Groq → OpenAI behavior for users who had both keys set as a true failover. This restores the prior semantics through the unified abstraction: - New getAiFallbackClient() in ai-provider.ts returns an OpenAI client (with the same timeout / maxRetries settings) when both GROQ_API_KEY and OPENAI_API_KEY are set AND no AI_BASE_URL override is in effect. An explicit AI_BASE_URL means the user has chosen a specific provider; no implicit fallback is added in that case. - callAiApi in workflows/generate-ai.ts wraps the primary call in a try/catch; on any primary failure, if a fallback client is available it retries once with OpenAI before propagating. JSON-mode handling is applied to both legs via a shared invokeChat helper. - A console.warn surfaces the fallback so the failure is observable.
|
Closing: the LLM half is superseded by #2113 (AI_PROVIDER=openai-compatible with AI_BASE_URL / AI_MODEL / AI_API_KEY), and the STT half targeted the Deepgram path, which main has since replaced with AssemblyAI. Local transcription is still missing on main, so I'll open a smaller follow-up that adds an OpenAI-compatible STT option (faster-whisper-server, whisper.cpp, etc.) alongside AssemblyAI, built on the new provider layer. |
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Summary
Depends on #1874. That proxy fix is the prerequisite for any workflow to execute on a self-host. Without #1874 merged, this PR's transcription path queues but never runs (the workflow runtime's /.well-known/workflow/v1/* HTTP callbacks get 307→/login). Once #1874 lands, this PR rebases cleanly onto main with no further changes.
Root cause
Self-hosted Cap currently requires three paid third-party providers (Deepgram + Groq/OpenAI) to make the share-page AI features (summary, chapters, transcript) work. Even with paid keys, the three calls use three inconsistent code paths:
@xenova in #1356: "any interest in using a local model for speech transcription? 👀"
PR #1705 already ships local STT in the desktop app (Parakeet). Local models are part of Cap's stance — just not on the server side, yet.
Fix
Collapse the three call patterns into one OpenAI-compatible client abstraction, configured by env. The OpenAI API is the lingua franca that Groq, OpenAI, Ollama, vLLM, OpenRouter, LiteLLM, faster-whisper-server, and whisper.cpp's HTTP server all already speak.
Two concerns → two env triples (all optional, all default to existing behavior):
Key simplification on the STT path: OpenAI's /v1/audio/transcriptions natively returns WebVTT (response_format: "vtt"). That's exactly what Cap already writes to S3, so when STT_BASE_URL is set the Deepgram-specific formatToWebVTT(DeepgramResult) adapter drops out — the rest of the pipeline is unchanged.
Commits are split into 4 logical groups, each typechecking independently:
End-state for a fully-local self-host (FYI — not part of this PR's required setup):
Backwards compatibility
If the new env vars are unset, behavior is identical to today: GROQ_API_KEY → Groq path; OPENAI_API_KEY → existing OpenAI fallback; DEEPGRAM_API_KEY → Deepgram. The Groq path now constructs an openai SDK client with baseURL = https://api.groq.com/openai/v1 — same wire protocol, no observable difference.
Verification
End-to-end test on a local Docker Compose self-host with Ollama (Gemma 3 12B) + hwdsl2/whisper-server (Whisper base), after applying both #1874 and this PR:
Before this PR (with #1874 applied alone — workflow runtime works)
After this PR
Gates clean for changed files: pnpm exec biome check --write, pnpm exec tsc -b, pnpm vitest run __tests__/unit/generate-ai-title.test.ts (6/6).
Out of scope (intentional)
Related
Design questions
Greptile Summary
This PR replaces three divergent AI call patterns (Groq SDK, raw OpenAI fetch, Deepgram SDK) with a single OpenAI-compatible client abstraction in lib/ai-provider.ts, gated by six new optional env vars (AI_BASE_URL/KEY/MODEL, STT_BASE_URL/KEY/MODEL). Existing deployments using GROQ_API_KEY / OPENAI_API_KEY / DEEPGRAM_API_KEY are unaffected by default.
Confidence Score: 3/5
Safe to merge for new self-hosted deployments; two unintended behavioral changes affect existing dual-key setups and the support chatbot routing.
Two issues need resolution before merge. First, workflows/generate-ai.ts silently drops the Groq→OpenAI failover — users with both keys set lose automatic recovery from Groq downtime. Second, lib/messenger/agent.ts was migrated despite the PR explicitly calling it out of scope, so the support chatbot's final fallback now routes through any AI_BASE_URL-configured provider (e.g. a local Ollama instance), which may produce unsuitable responses for a customer support context.
apps/web/workflows/generate-ai.ts (failover removal) and apps/web/lib/messenger/agent.ts (unintended migration of the support chatbot)
Important Files Changed
Comments Outside Diff (2)
-
Messenger agent was migrated — contradicts PR description
-
Double-cast through unknown to string
Prompt To Fix All With AIapps/web/lib/messenger/agent.ts, line 168-191 (link)
The PR's "Out of scope" section explicitly states: "The Anthropic + OpenAI raw-fetch fallbacks in apps/web/lib/messenger/agent.ts were NOT migrated." But callGroq has been renamed to callAiProvider and now calls getAiClient(). For any self-hosted deployment that sets AI_BASE_URL (the stated target of this PR) but has neither ANTHROPIC_API_KEY nor OPENAI_API_KEY, the support chatbot's last resort will now be a local Ollama/vLLM instance. A local Gemma model answering customer support queries is likely not the intended behavior, and it contradicts the stated out-of-scope decision.
Prompt To Fix With AIapps/web/workflows/transcribe.ts, line 536-541 (link)
The OpenAI Node SDK v4 overloads audio.transcriptions.create — when response_format is "vtt", "srt", or "text" the runtime value is a plain string, but the TypeScript generic signature falls through to the Transcription type. The as unknown as string workaround is valid here, but leaving a comment explaining why the cast is needed would help the next reader avoid accidentally "fixing" it by removing the cast (which would cause a type error on .includes("WEBVTT")).
Prompt To Fix With AINote: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Reviews (1): Last reviewed commit: "chore(docker): expose AI/STT env vars th..." | Re-trigger Greptile