mirror of
https://github.com/tiennm99/goclaw.git
synced 2026-10-11 03:13:24 +00:00
* fix(agent): filter skill slash commands by the agent's grants The inline <available_skills> block was already filtered by visibility and agent grants, but slash activation read the loader's full skill list. A /<slug> command could therefore activate a skill the agent was never granted, and the not-found suggestions disclosed that such skills existed. Resolve slash commands against the agent's allow list via FilterSkills. The list is filtered before matching, so /list-skills, /help and the near-match suggestions are all covered by the same change. The allow-list convention is unchanged: nil means every skill (the fallback when the access store errors), an empty slice means none, and a populated slice is an explicit set of slugs. * fix(agent): split slash commands on any whitespace The tokenizer separated the skill name from the rest of the message with strings.Cut(after, " ") — a literal space. A user who typed the command and pressed Enter before the rest of the message sent "/ck:git\nreview the diff", which parsed to the target "ck:git\nreview" and matched no skill. The reply was "skill not found" followed by near-matches that included the skill they had just named, because the similarity fallback searched the mangled string and still landed beside it. Multi-line messages are ordinary in Slack and Telegram, so this was reachable in normal use. Introduce cutFirstField, which splits on the first run of whitespace, and use it everywhere the literal-space split appeared. The match loop had the same assumption in strings.HasPrefix(raw, value+" "); hasFieldPrefix replaces it and decodes the following rune properly, so a multi-byte skill name is not truncated. That also stops a plain string prefix from matching: a command for "frontend-design-extra" no longer resolves to "frontend-design". * fix(tools): inherit the parent agent's tool policy in spawned subagents buildSubagentToolsRegistry clones the parent registry, then overwrites exec, read_file, write_file and list_files with freshly constructed tools. A fresh tool carries none of the hardening the gateway applies to the parent's instances at startup: exec path denials and their exemptions, the shell deny-group toggles, the command keyword allowlist, and the read/write/list deny prefixes covering config.json, the internal databases and delegate/. Spawning a subagent therefore widened what an agent could reach — the parent was blocked from the data dir, the subagent was not. Verified by disabling the new call: the subagent's exec read config.json out of the denied data dir, read_file did the same, and write_file overwrote it. Copy the policy from the live parent instances rather than re-deriving it, so there is one source of truth and a deny path reloaded later through config pub/sub reaches subagents without a second wiring site that can drift. Allow-prefixes are inherited alongside the denials: they are what make the denied roots usable at all, since the skills store sits under the denied data dir. Copying denials without them would leave a subagent unable to read the skills it is told to use. * fix(agent): use one inline-vs-search decision for prompt and preview The system-prompt preview exists to show the prompt an agent actually gets, but it decided between inline skills and search mode on its own terms: it counted tokens with the fallback counter over the fully rendered XML — tags and <location> paths included, at roughly runes/2 — while the prompt builder estimated name+description at chars/4. On the same eighteen skills that read 3087 against 944, so the preview reported search mode for an agent that was running inline. Extract shouldInlineSkills and call it from both paths. The estimate keeps the prompt builder's rule, mirroring BuildSummary's 200-rune description truncation, so the decision still costs no rendering. The preview's loader interface gains FilterSkills to feed it. Its behaviour on an access-store error is unchanged: an empty allow list yields no skills rather than falling back to showing every skill. * fix(http): keep reference files when importing skills Export archives the whole skill directory, but import recognised only metadata.json, SKILL.md and grants.jsonl. The switch had no default, so every other file was discarded without a log line. A skill whose SKILL.md cites references/*.md arrived without them, the import reported success, and the skill failed at the first read. Collect the remaining entries and write them under the skill directory with their structure intact. Archive entry names are attacker-controlled, so each path goes through sanitizeRelPath: a traversal attempt collapses to a relative path and the write stays inside the skill directory. * test(agent): make the inline-decision table exercise both gates The "100 skills, 200-char descriptions" case claimed to cover the token ceiling, but 100 exceeds the count ceiling so the count check short-circuited and the token branch never ran — deleting that branch would not have failed the test. Split the table so each case isolates one gate: tiny descriptions for the count rows, a count safely under the ceiling for the token rows. Removing the token comparison now fails the over-the-ceiling case and nothing else. * fix(agent): gate slash commands on the managed tier only The allow list comes from a query over the `skills` table, so filesystem-tier skills — the workspace, .agents and ~/.agents directories of the five-tier loader — have no row in it. Filtering every skill against the list made those four tiers unreachable by slash while skill_search still found them unfiltered, which turned a security fix into a functional regression. Apply the list to managed skills only. Builtins are seeded with is_system and returned unconditionally, so they stay reachable either way. This closes slash activation of ungranted managed skills. It does not close skill_search and use_skill, which still read the loader's full list; that is a wider change and wants its own review. * fix(http): stop imported skill files from replacing guarded ones Two ways the auxiliary write could go wrong. GuardSkillContent scans SKILL.md before anything reaches disk, but the switch that routes archive entries compares the raw path. An entry named "./SKILL.md" does not match it, so it fell through to the auxiliary set — where sanitizeRelPath collapses it back to "SKILL.md" and the write replaced the file that had just been scanned. Reject any auxiliary path that resolves to a name the import handles by itself, and log the attempt. Tar directory entries carry a trailing separator and no content. Written as files they took the name a real directory needed, so everything beneath was dropped — and whether that happened depended on map iteration order, making it intermittent. Skip them; MkdirAll creates what is needed. Also cap the number of auxiliary files per skill and log when the cap trims an archive, so a truncated import is visible rather than silent. --------- Co-authored-by: mor-phongdt <phong.dangtuan@mor.com.vn>
477 lines
17 KiB
Go
477 lines
17 KiB
Go
package cmd
|
|
|
|
import (
|
|
"context"
|
|
"log/slog"
|
|
|
|
"github.com/nextlevelbuilder/goclaw/internal/audio/elevenlabs"
|
|
geminiaudio "github.com/nextlevelbuilder/goclaw/internal/audio/gemini"
|
|
minimaxaudio "github.com/nextlevelbuilder/goclaw/internal/audio/minimax"
|
|
"github.com/nextlevelbuilder/goclaw/internal/audio/openaicompat"
|
|
"github.com/nextlevelbuilder/goclaw/internal/bus"
|
|
"github.com/nextlevelbuilder/goclaw/internal/config"
|
|
"github.com/nextlevelbuilder/goclaw/internal/memory"
|
|
"github.com/nextlevelbuilder/goclaw/internal/orchestration"
|
|
"github.com/nextlevelbuilder/goclaw/internal/providers"
|
|
"github.com/nextlevelbuilder/goclaw/internal/sandbox"
|
|
"github.com/nextlevelbuilder/goclaw/internal/store"
|
|
"github.com/nextlevelbuilder/goclaw/internal/tools"
|
|
"github.com/nextlevelbuilder/goclaw/internal/tts"
|
|
usagecaps "github.com/nextlevelbuilder/goclaw/internal/usage/caps"
|
|
)
|
|
|
|
// resolveEmbeddingProvider selects an embedding provider from DB only.
|
|
// Resolution order:
|
|
// 1. system_configs "embedding.provider" + "embedding.model" → find provider by name in DB
|
|
// 2. Auto-detect: first DB provider with settings.embedding.enabled = true
|
|
//
|
|
// Config file (agents.defaults.memory) is NOT used — all embedding config lives in DB.
|
|
func resolveEmbeddingProvider(
|
|
providerStore store.ProviderStore,
|
|
providerReg *providers.Registry,
|
|
sysConfigs store.SystemConfigStore,
|
|
) memory.EmbeddingProvider {
|
|
masterCtx := store.WithTenantID(context.Background(), store.MasterTenantID) // for system_configs (tenant-scoped)
|
|
|
|
// 1. System config: embedding.provider (set via UI / API)
|
|
if sysConfigs != nil {
|
|
if name, err := sysConfigs.Get(masterCtx, "embedding.provider"); err == nil && name != "" {
|
|
sysModel := ""
|
|
if m, mErr := sysConfigs.Get(masterCtx, "embedding.model"); mErr == nil {
|
|
sysModel = m
|
|
}
|
|
var mcfg *config.MemoryConfig
|
|
if sysModel != "" {
|
|
mcfg = &config.MemoryConfig{EmbeddingModel: sysModel}
|
|
}
|
|
p := resolveEmbeddingFromDB(masterCtx, providerStore, name, mcfg, providerReg)
|
|
if p != nil {
|
|
slog.Info("embedding provider from system_configs", "name", name, "model", p.Model())
|
|
return p
|
|
}
|
|
slog.Warn("system_configs embedding.provider not found in DB", "name", name)
|
|
}
|
|
}
|
|
|
|
// 2. Auto-detect only from the master tenant. The selected provider receives
|
|
// cross-tenant semantic content during startup maintenance, so a tenant-owned
|
|
// endpoint must never be selected for this process-wide responsibility.
|
|
allProviders, err := providerStore.ListProviders(masterCtx)
|
|
if err != nil {
|
|
slog.Warn("failed to list providers for embedding auto-detect", "error", err)
|
|
return nil
|
|
}
|
|
|
|
for _, dbp := range allProviders {
|
|
if !dbp.Enabled || store.NoEmbeddingTypes[dbp.ProviderType] {
|
|
continue
|
|
}
|
|
es := store.ParseEmbeddingSettings(dbp.Settings)
|
|
if es == nil || !es.Enabled {
|
|
continue
|
|
}
|
|
|
|
ep := buildEmbeddingProvider(&dbp, es, nil, providerReg)
|
|
if ep != nil {
|
|
slog.Info("embedding provider auto-detected", "name", dbp.Name, "model", ep.Model())
|
|
return ep
|
|
}
|
|
}
|
|
|
|
slog.Warn("no embedding provider configured — enable embedding in a provider's settings")
|
|
return nil
|
|
}
|
|
|
|
// resolveEmbeddingFromDB finds a specific provider by name and builds an embedding provider from it.
|
|
func resolveEmbeddingFromDB(
|
|
ctx context.Context,
|
|
providerStore store.ProviderStore,
|
|
name string,
|
|
memCfg *config.MemoryConfig,
|
|
providerReg *providers.Registry,
|
|
) memory.EmbeddingProvider {
|
|
dbp, err := providerStore.GetProviderByName(ctx, name)
|
|
if err != nil {
|
|
return nil
|
|
}
|
|
if !dbp.Enabled {
|
|
slog.Warn("embedding provider disabled", "name", name)
|
|
return nil
|
|
}
|
|
if store.NoEmbeddingTypes[dbp.ProviderType] {
|
|
slog.Warn("embedding provider type does not support embeddings", "name", name, "type", dbp.ProviderType)
|
|
return nil
|
|
}
|
|
es := store.ParseEmbeddingSettings(dbp.Settings)
|
|
return buildEmbeddingProvider(dbp, es, memCfg, providerReg)
|
|
}
|
|
|
|
// buildEmbeddingProvider creates a memory.EmbeddingProvider from a DB provider record.
|
|
func buildEmbeddingProvider(
|
|
dbp *store.LLMProviderData,
|
|
es *store.EmbeddingSettings,
|
|
memCfg *config.MemoryConfig,
|
|
providerReg *providers.Registry,
|
|
) memory.EmbeddingProvider {
|
|
// Resolve model: embedding settings → memCfg override → default
|
|
model := "text-embedding-3-small"
|
|
if es != nil && es.Model != "" {
|
|
model = es.Model
|
|
}
|
|
if memCfg != nil && memCfg.EmbeddingModel != "" {
|
|
model = memCfg.EmbeddingModel
|
|
}
|
|
|
|
// Resolve API base: embedding settings → provider api_base → registry
|
|
apiBase := dbp.APIBase
|
|
if es != nil && es.APIBase != "" {
|
|
apiBase = es.APIBase
|
|
}
|
|
if memCfg != nil && memCfg.EmbeddingAPIBase != "" {
|
|
apiBase = memCfg.EmbeddingAPIBase
|
|
}
|
|
|
|
// Dimension truncation: default to RequiredMemoryEmbeddingDimensions to match pgvector schema.
|
|
// Models that natively output 1536 ignore the parameter; models with larger native dims get truncated.
|
|
dims := store.RequiredMemoryEmbeddingDimensions
|
|
if es != nil && es.Dimensions > 0 && es.Dimensions != store.RequiredMemoryEmbeddingDimensions {
|
|
slog.Warn("ignoring incompatible provider embedding dimensions for memory schema",
|
|
"provider", dbp.Name, "requested", es.Dimensions, "required", store.RequiredMemoryEmbeddingDimensions)
|
|
}
|
|
|
|
// Try registry first for the actual API key / base (handles runtime-registered providers)
|
|
if providerReg != nil {
|
|
if regProv, regErr := providerReg.Get(context.Background(), dbp.Name); regErr == nil {
|
|
if op, ok := regProv.(*providers.OpenAIProvider); ok {
|
|
if apiBase == "" {
|
|
apiBase = op.APIBase()
|
|
}
|
|
ep := memory.NewOpenAIEmbeddingProvider(dbp.Name, op.APIKey(), apiBase, model)
|
|
ep.WithDimensions(dims)
|
|
return ep
|
|
}
|
|
slog.Debug("embedding provider in registry is not OpenAI-compatible, using DB record", "name", dbp.Name)
|
|
}
|
|
}
|
|
|
|
// Fallback: build directly from DB record
|
|
if dbp.APIKey != "" {
|
|
ep := memory.NewOpenAIEmbeddingProvider(dbp.Name, dbp.APIKey, apiBase, model)
|
|
ep.WithDimensions(dims)
|
|
return ep
|
|
}
|
|
|
|
return nil
|
|
}
|
|
|
|
func setupSubagents(providerReg *providers.Registry, cfg *config.Config, msgBus *bus.MessageBus, toolsReg *tools.Registry, workspace string, sandboxMgr sandbox.Manager, secureCLIStore store.SecureCLIStore, usageCapSvc *usagecaps.Service, admission *orchestration.ChildRunAdmission) *tools.SubagentManager {
|
|
names := providerReg.List(context.Background())
|
|
if len(names) == 0 {
|
|
return nil
|
|
}
|
|
|
|
agentCfg := cfg.ResolveAgent("default")
|
|
provider, err := providerReg.Get(context.Background(), agentCfg.Provider)
|
|
if err != nil {
|
|
provider, _ = providerReg.Get(context.Background(), names[0])
|
|
}
|
|
if provider == nil {
|
|
return nil
|
|
}
|
|
|
|
subCfg := tools.DefaultSubagentConfig()
|
|
|
|
// Apply config file overrides if present (matching TS agents.defaults.subagents).
|
|
if sc := agentCfg.Subagents; sc != nil {
|
|
if sc.MaxConcurrent > 0 {
|
|
subCfg.MaxConcurrent = sc.MaxConcurrent
|
|
}
|
|
if sc.MaxSpawnDepth > 0 {
|
|
subCfg.MaxSpawnDepth = min(sc.MaxSpawnDepth, 5) // TS: max 5
|
|
}
|
|
if sc.MaxChildrenPerAgent > 0 {
|
|
subCfg.MaxChildrenPerAgent = min(sc.MaxChildrenPerAgent, 20) // TS: max 20
|
|
}
|
|
if sc.ArchiveAfterMinutes > 0 {
|
|
subCfg.ArchiveAfterMinutes = sc.ArchiveAfterMinutes
|
|
}
|
|
if sc.MaxRetries > 0 {
|
|
subCfg.MaxRetries = sc.MaxRetries
|
|
}
|
|
if sc.Model != "" {
|
|
subCfg.Model = sc.Model
|
|
}
|
|
}
|
|
|
|
// Tool factory: clone parent registry (inherits web_fetch, web_search, browser, MCP tools, etc.)
|
|
// then override file/exec tools with workspace-scoped versions.
|
|
// NOTE: SubagentManager.applyDenyList() handles deny lists after createTools(),
|
|
// so we don't apply deny lists here.
|
|
toolsFactory := func() *tools.Registry {
|
|
reg, _ := buildSubagentToolsRegistry(toolsReg, workspace, agentCfg.RestrictToWorkspace, sandboxMgr, secureCLIStore)
|
|
return reg
|
|
}
|
|
|
|
manager := tools.NewSubagentManagerWithAdmission(provider, providerReg, agentCfg.Model, msgBus, toolsFactory, subCfg, admission)
|
|
manager.SetUsageCapService(usageCapSvc)
|
|
manager.SetAgentBudget(agentCfg.ContextWindow, agentCfg.MaxTokens)
|
|
return manager
|
|
}
|
|
|
|
// buildSubagentToolsRegistry produces a cloned tool registry for a subagent
|
|
// with workspace-scoped file/exec tools registered. The returned ExecTool is
|
|
// also returned for test assertion (Red Team F3): callers verify that the
|
|
// secureCLIStore is wired so the subagent's exec path enforces the gate.
|
|
func buildSubagentToolsRegistry(
|
|
parentReg *tools.Registry,
|
|
workspace string,
|
|
restrict bool,
|
|
sandboxMgr sandbox.Manager,
|
|
secureCLIStore store.SecureCLIStore,
|
|
) (*tools.Registry, *tools.ExecTool) {
|
|
reg := parentReg.Clone()
|
|
var execTool *tools.ExecTool
|
|
var readTool *tools.ReadFileTool
|
|
var writeTool *tools.WriteFileTool
|
|
var listTool *tools.ListFilesTool
|
|
if sandboxMgr != nil {
|
|
readTool = tools.NewSandboxedReadFileTool(workspace, restrict, sandboxMgr)
|
|
writeTool = tools.NewSandboxedWriteFileTool(workspace, restrict, sandboxMgr)
|
|
listTool = tools.NewSandboxedListFilesTool(workspace, restrict, sandboxMgr)
|
|
execTool = tools.NewSandboxedExecTool(workspace, restrict, sandboxMgr)
|
|
} else {
|
|
readTool = tools.NewReadFileTool(workspace, restrict)
|
|
writeTool = tools.NewWriteFileTool(workspace, restrict)
|
|
listTool = tools.NewListFilesTool(workspace, restrict)
|
|
execTool = tools.NewExecTool(workspace, restrict)
|
|
}
|
|
|
|
// These four tools are built fresh, so they start with none of the hardening the
|
|
// gateway applied to the parent's instances at startup: exec path denials and their
|
|
// exemptions, shell deny-group toggles, the command keyword allowlist, and the
|
|
// read/write/list deny prefixes covering config.json, the databases, and delegate/.
|
|
// Without this, spawning a subagent widened reach — the parent could not touch the
|
|
// data dir, the subagent could. Inherit from the live parent instances so there is
|
|
// one source of truth and a later config reload cannot leave subagents behind.
|
|
inheritParentPathPolicy(parentReg, readTool, writeTool, listTool, execTool)
|
|
|
|
reg.Register(readTool)
|
|
reg.Register(writeTool)
|
|
reg.Register(listTool)
|
|
reg.Register(execTool)
|
|
// Red Team F3: subagent ExecTool must enforce the secure-CLI gate
|
|
// (and env scrub on fall-through) — without this, a parent agent
|
|
// can spawn a subagent to bypass the gate via host-inherited env.
|
|
if secureCLIStore != nil {
|
|
execTool.SetSecureCLIStore(secureCLIStore)
|
|
}
|
|
return reg, execTool
|
|
}
|
|
|
|
// inheritParentPathPolicy copies the parent registry's tool hardening onto the freshly
|
|
// built subagent tools. A tool missing from the parent registry, or registered there
|
|
// under an unexpected concrete type, is skipped: the subagent then has no policy to
|
|
// inherit for it, which matches the parent having none to give.
|
|
func inheritParentPathPolicy(
|
|
parentReg *tools.Registry,
|
|
readTool *tools.ReadFileTool,
|
|
writeTool *tools.WriteFileTool,
|
|
listTool *tools.ListFilesTool,
|
|
execTool *tools.ExecTool,
|
|
) {
|
|
if parentReg == nil {
|
|
return
|
|
}
|
|
if pt, ok := parentReg.Get("read_file"); ok {
|
|
if parent, ok := pt.(*tools.ReadFileTool); ok {
|
|
readTool.InheritPathPolicy(parent)
|
|
}
|
|
}
|
|
if pt, ok := parentReg.Get("write_file"); ok {
|
|
if parent, ok := pt.(*tools.WriteFileTool); ok {
|
|
writeTool.InheritPathPolicy(parent)
|
|
}
|
|
}
|
|
if pt, ok := parentReg.Get("list_files"); ok {
|
|
if parent, ok := pt.(*tools.ListFilesTool); ok {
|
|
listTool.InheritPathPolicy(parent)
|
|
}
|
|
}
|
|
if pt, ok := parentReg.Get("exec"); ok {
|
|
if parent, ok := pt.(*tools.ExecTool); ok {
|
|
execTool.InheritSecurityPolicy(parent)
|
|
}
|
|
}
|
|
}
|
|
|
|
// setupTTS creates the TTS manager from config and registers providers.
|
|
// Edge TTS is always registered (free, no API key required).
|
|
// Always returns a non-nil manager with at least one provider.
|
|
func setupTTS(cfg *config.Config) *tts.Manager {
|
|
ttsCfg := cfg.Tts
|
|
|
|
mgr := tts.NewManager(tts.ManagerConfig{
|
|
Primary: ttsCfg.Provider,
|
|
Auto: tts.AutoMode(ttsCfg.Auto),
|
|
Mode: tts.Mode(ttsCfg.Mode),
|
|
MaxLength: ttsCfg.MaxLength,
|
|
TimeoutMs: ttsCfg.TimeoutMs,
|
|
})
|
|
|
|
// Register providers that have API keys configured
|
|
if key := ttsCfg.OpenAI.APIKey; key != "" {
|
|
mgr.RegisterProvider(tts.NewOpenAIProvider(tts.OpenAIConfig{
|
|
APIKey: key,
|
|
APIBase: ttsCfg.OpenAI.APIBase,
|
|
Model: ttsCfg.OpenAI.Model,
|
|
Voice: ttsCfg.OpenAI.Voice,
|
|
TimeoutMs: ttsCfg.TimeoutMs,
|
|
}))
|
|
}
|
|
|
|
// OpenAI-compatible self-hosted endpoint. api_base is the enable switch —
|
|
// unlike the vendor providers there is no API key to gate on, since these
|
|
// endpoints commonly have no auth.
|
|
if base := ttsCfg.OpenAICompat.APIBase; base != "" {
|
|
provider, err := openaicompat.NewTTSProvider(openaicompat.Config{
|
|
APIBase: base,
|
|
APIKey: ttsCfg.OpenAICompat.APIKey,
|
|
TTSModel: ttsCfg.OpenAICompat.Model,
|
|
TTSVoice: ttsCfg.OpenAICompat.Voice,
|
|
TTSFormat: ttsCfg.OpenAICompat.Format,
|
|
TimeoutMs: ttsCfg.TimeoutMs,
|
|
})
|
|
if err != nil {
|
|
slog.Warn("audio.tts: openai_compat not registered", "error", err)
|
|
} else {
|
|
mgr.RegisterProvider(provider)
|
|
slog.Info("audio.tts: openai_compat registered", "api_base", base)
|
|
}
|
|
}
|
|
|
|
if key := ttsCfg.ElevenLabs.APIKey; key != "" {
|
|
mgr.RegisterProvider(tts.NewElevenLabsProvider(tts.ElevenLabsConfig{
|
|
APIKey: key,
|
|
BaseURL: ttsCfg.ElevenLabs.BaseURL,
|
|
VoiceID: ttsCfg.ElevenLabs.VoiceID,
|
|
ModelID: ttsCfg.ElevenLabs.ModelID,
|
|
TimeoutMs: ttsCfg.TimeoutMs,
|
|
}))
|
|
}
|
|
|
|
// Edge TTS is free (no API key) — always register so it's available as primary or fallback.
|
|
mgr.RegisterProvider(tts.NewEdgeProvider(tts.EdgeConfig{
|
|
Voice: ttsCfg.Edge.Voice,
|
|
Rate: ttsCfg.Edge.Rate,
|
|
TimeoutMs: ttsCfg.TimeoutMs,
|
|
}))
|
|
|
|
if key := ttsCfg.MiniMax.APIKey; key != "" {
|
|
mgr.RegisterProvider(tts.NewMiniMaxProvider(tts.MiniMaxConfig{
|
|
APIKey: key,
|
|
GroupID: ttsCfg.MiniMax.GroupID,
|
|
APIBase: ttsCfg.MiniMax.APIBase,
|
|
Model: ttsCfg.MiniMax.Model,
|
|
VoiceID: ttsCfg.MiniMax.VoiceID,
|
|
TimeoutMs: ttsCfg.TimeoutMs,
|
|
}))
|
|
}
|
|
|
|
if key := ttsCfg.Gemini.APIKey; key != "" {
|
|
mgr.RegisterProvider(geminiaudio.NewProvider(geminiaudio.Config{
|
|
APIKey: key,
|
|
APIBase: ttsCfg.Gemini.APIBase,
|
|
Voice: ttsCfg.Gemini.Voice,
|
|
Model: ttsCfg.Gemini.Model,
|
|
TimeoutMs: ttsCfg.TimeoutMs,
|
|
}))
|
|
}
|
|
|
|
if !mgr.HasProviders() {
|
|
return nil
|
|
}
|
|
|
|
return mgr
|
|
}
|
|
|
|
// setupAudioExtras wires Music and SFX providers into the audio Manager.
|
|
// ElevenLabs is registered for both SFX and Music when an API key is present.
|
|
// MiniMax music is registered when cfg.Audio.Music is configured with a key.
|
|
// Phase 4 will add STT providers here.
|
|
func setupAudioExtras(cfg *config.Config, mgr *tts.Manager) {
|
|
ellKey := cfg.Tts.ElevenLabs.APIKey
|
|
ellBase := cfg.Tts.ElevenLabs.BaseURL
|
|
|
|
// ElevenLabs SFX — reuse TTS credentials.
|
|
if ellKey != "" {
|
|
mgr.RegisterSFX(elevenlabs.NewSFXProvider(elevenlabs.Config{
|
|
APIKey: ellKey,
|
|
BaseURL: ellBase,
|
|
}))
|
|
slog.Info("audio.sfx: elevenlabs registered")
|
|
}
|
|
|
|
// ElevenLabs Music — same credentials, uses /v1/music endpoint.
|
|
if ellKey != "" {
|
|
mgr.RegisterMusic(elevenlabs.NewMusicProvider(elevenlabs.Config{
|
|
APIKey: ellKey,
|
|
BaseURL: ellBase,
|
|
}))
|
|
slog.Info("audio.music: elevenlabs registered")
|
|
}
|
|
|
|
// MiniMax Music — optional, from cfg.Audio.Music block.
|
|
if cfg.Audio != nil && cfg.Audio.Music != nil {
|
|
mc := cfg.Audio.Music
|
|
if mc.APIKey != "" {
|
|
mgr.RegisterMusic(minimaxaudio.NewMusicProvider(minimaxaudio.MusicConfig{
|
|
APIKey: mc.APIKey,
|
|
APIBase: mc.BaseURL,
|
|
Model: mc.Model,
|
|
}))
|
|
slog.Info("audio.music: minimax registered")
|
|
}
|
|
}
|
|
|
|
// STT chain. Built from what actually registered, so Transcribe never walks
|
|
// a name it will only skip with a warning. "proxy" is always appended: it is
|
|
// registered later, per channel, by BridgeLegacySTT.
|
|
var sttChain []string
|
|
|
|
// OpenAI-compatible self-hosted endpoint, first when present: it is an
|
|
// explicit operator choice, and it keeps audio on the local network.
|
|
if base := cfg.Tts.OpenAICompat.APIBase; base != "" {
|
|
provider, err := openaicompat.NewSTTProvider(openaicompat.Config{
|
|
APIBase: base,
|
|
APIKey: cfg.Tts.OpenAICompat.APIKey,
|
|
STTModel: cfg.Tts.OpenAICompat.STTModel,
|
|
TimeoutMs: cfg.Tts.TimeoutMs,
|
|
})
|
|
if err != nil {
|
|
slog.Warn("audio.stt: openai_compat not registered", "error", err)
|
|
} else {
|
|
mgr.RegisterSTT(provider)
|
|
sttChain = append(sttChain, provider.Name())
|
|
slog.Info("audio.stt: openai_compat registered", "api_base", base)
|
|
}
|
|
}
|
|
|
|
// ElevenLabs STT (Scribe v2) — reuse TTS credentials. Registered as tenant-scope
|
|
// default; per-request tenant override lands via builtin_tools[stt] in Phase 5
|
|
// channel migration. Legacy per-channel STTProxyURL is bridged separately.
|
|
if ellKey != "" {
|
|
mgr.RegisterSTT(elevenlabs.NewSTTProvider(elevenlabs.Config{
|
|
APIKey: ellKey,
|
|
BaseURL: ellBase,
|
|
}))
|
|
sttChain = append(sttChain, "elevenlabs")
|
|
slog.Info("audio.stt: elevenlabs registered")
|
|
}
|
|
|
|
if len(sttChain) > 0 {
|
|
sttChain = append(sttChain, "proxy")
|
|
mgr.SetSTTChain(sttChain)
|
|
slog.Info("audio.stt: chain configured", "chain", sttChain)
|
|
}
|
|
}
|