2026-06-21 - 2026-09-21

Overview

14 Active Pull Requests
40 Active Issues
Excluding merges, 2 authors have pushed 70 commits to main and 101 commits to all branches. On main, 53 files have changed and there have been 8675 additions and 3831 deletions.

14 Pull requests merged by 1 user

31 Issues closed from 1 user

Closed #5 Verify the stack on the real Radeon R9700 box (GPU passthrough + Lazytainer idle-stop) 2026-09-06 20:30:07 +00:00

Closed #43 scripts/switch-model.sh: swap GPU residency between Qwen and ComfyUI 2026-09-06 19:01:03 +00:00

Closed #42 Extend downloader for the diffusion model weights 2026-09-06 19:01:02 +00:00

Closed #44 Add llama-server-fast: small non-thinking classifier/fast model for qwen-code Auto Mode 2026-09-06 18:39:31 +00:00

Closed #39 Best diffusion model given full ~32GB headroom (Qwen stopped) 2026-09-05 19:16:32 +00:00

Closed #40 Why lazytainer idle-stop doesn't work today, and how the switch script should handle it 2026-09-05 18:57:39 +00:00

Closed #37 OmniRoute: docker-compose service design and secrets/config plan 2026-09-03 17:38:46 +00:00

Closed #32 OmniRoute: can it route to arbitrary OpenAI-compatible local endpoints (llama-server), not just Ollama? 2026-09-03 17:24:11 +00:00

Closed #36 OmniRoute: per-workload virtual keys, admin UI, and deployment shape (Docker/compose) 2026-09-03 17:21:22 +00:00

Closed #35 OmniRoute: web-search tool equivalent to the SearXNG standalone endpoint 2026-09-03 17:21:16 +00:00

Closed #33 OmniRoute due-diligence: maintainer, repo history, npm package trust 2026-09-03 17:20:59 +00:00

Closed #34 OmniRoute: memory/knowledgebase parity with litellm-pgvector 2026-09-03 17:20:50 +00:00

Closed #29 Point opencode's provider block at litellm.home with a virtual key 2026-09-03 04:42:01 +00:00

Closed #28 Should opencode's context/output limits be generated from litellm-config.yaml, or hand-maintained? 2026-09-03 04:42:00 +00:00

Closed #27 How does opencode's auto-compact actually work — trigger threshold and config surface? 2026-09-03 04:26:08 +00:00

Closed #25 LangChain+pgvector direct RAG vs. the litellm-pgvector connector — which fits this stack? 2026-09-02 20:09:50 +00:00

Closed #23 How does LiteLLM's knowledgebase/vector-store feature work, and what does it need? 2026-09-02 19:39:30 +00:00

Closed #22 How does LiteLLM's SearXNG web-search integration work, and what does wiring it in require? 2026-09-02 19:38:40 +00:00

Closed #15 Migrate Open WebUI (and existing consumers) to route through the new proxy 2026-08-25 05:18:04 +00:00

Closed #14 Author the docker-compose service for the chosen proxy 2026-08-25 05:11:17 +00:00

Closed #16 How should the proxy queue and prioritize concurrent requests across workloads? 2026-08-25 05:07:21 +00:00

Closed #13 Network/hostname plan for exposing the proxy (LAN + external) 2026-08-25 05:05:02 +00:00

Closed #12 How should per-workload API keys/accounts be provisioned, rotated, and documented? 2026-08-25 04:56:53 +00:00

Closed #11 Reference cloud model/pricing for the shadow-cost estimate 2026-08-25 04:49:50 +00:00

Closed #10 Which self-hosted AI gateway/proxy tool fits this effort's needs? 2026-08-25 04:45:24 +00:00

Closed #8 How is ai.home / ai.haylan.ch reverse-proxied and TLS-terminated? 2026-08-24 16:28:23 +00:00

Closed #6 Write local-usage docs for Claude Code CLI, Kimi CLI, and OpenCode CLI against the local endpoint 2026-08-24 16:20:05 +00:00

Closed #7 Research OpenCode CLI installation and local-endpoint setup 2026-08-24 16:17:37 +00:00

Closed #4 Author the docker-compose stack (llama.cpp + Open WebUI + Qdrant + Lazytainer) 2026-08-24 11:01:21 +00:00

Closed #2 Confirm quantization availability for Qwen3.8-27B 2026-08-24 10:17:03 +00:00

Closed #3 Confirm Qwen3.8-27B tool-calling compatibility with llama.cpp's Anthropic Messages API shim 2026-08-24 10:16:43 +00:00

40 Issues created by 1 user

Opened #1 Local AI inference stack: llama.cpp + Qwen3.8-27B on Radeon R9700 2026-08-24 10:09:29 +00:00

Opened #2 Confirm quantization availability for Qwen3.8-27B 2026-08-24 10:09:41 +00:00

Opened #3 Confirm Qwen3.8-27B tool-calling compatibility with llama.cpp's Anthropic Messages API shim 2026-08-24 10:09:42 +00:00

Opened #4 Author the docker-compose stack (llama.cpp + Open WebUI + Qdrant + Lazytainer) 2026-08-24 10:09:57 +00:00

Opened #5 Verify the stack on the real Radeon R9700 box (GPU passthrough + Lazytainer idle-stop) 2026-08-24 10:09:57 +00:00

Opened #6 Write local-usage docs for Claude Code CLI, Kimi CLI, and OpenCode CLI against the local endpoint 2026-08-24 10:09:58 +00:00

Opened #7 Research OpenCode CLI installation and local-endpoint setup 2026-08-24 16:13:20 +00:00

Opened #8 How is ai.home / ai.haylan.ch reverse-proxied and TLS-terminated? 2026-08-24 16:24:58 +00:00

Opened #9 AI gateway/proxy: routing, per-workload keys, and cloud-cost estimation for the local LLM stack 2026-08-25 04:39:30 +00:00

Opened #10 Which self-hosted AI gateway/proxy tool fits this effort's needs? 2026-08-25 04:39:41 +00:00

Opened #11 Reference cloud model/pricing for the shadow-cost estimate 2026-08-25 04:39:54 +00:00

Opened #12 How should per-workload API keys/accounts be provisioned, rotated, and documented? 2026-08-25 04:39:55 +00:00

Opened #13 Network/hostname plan for exposing the proxy (LAN + external) 2026-08-25 04:39:56 +00:00

Opened #14 Author the docker-compose service for the chosen proxy 2026-08-25 04:40:05 +00:00

Opened #15 Migrate Open WebUI (and existing consumers) to route through the new proxy 2026-08-25 04:40:06 +00:00

Opened #16 How should the proxy queue and prioritize concurrent requests across workloads? 2026-08-25 04:41:13 +00:00

Opened #17 Verify the AI proxy stack on the real R9700 box (LiteLLM scheduler smoke test) 2026-08-25 05:11:03 +00:00

Opened #21 LiteLLM gateway: search, knowledgebase, and gateway-level memory 2026-09-02 19:18:43 +00:00

Opened #22 How does LiteLLM's SearXNG web-search integration work, and what does wiring it in require? 2026-09-02 19:18:54 +00:00

Opened #23 How does LiteLLM's knowledgebase/vector-store feature work, and what does it need? 2026-09-02 19:18:55 +00:00

Opened #24 Verify search + knowledgebase wiring (#21) on the real R9700 box 2026-09-02 19:37:43 +00:00

Opened #25 LangChain+pgvector direct RAG vs. the litellm-pgvector connector — which fits this stack? 2026-09-02 19:55:46 +00:00

Opened #26 opencode: context size and compaction settings for the litellm/qwen3.8-27b-local stack 2026-09-03 04:19:22 +00:00

Opened #27 How does opencode's auto-compact actually work — trigger threshold and config surface? 2026-09-03 04:19:33 +00:00

Opened #28 Should opencode's context/output limits be generated from litellm-config.yaml, or hand-maintained? 2026-09-03 04:19:45 +00:00

Opened #29 Point opencode's provider block at litellm.home with a virtual key 2026-09-03 04:19:56 +00:00

Opened #31 Migrate the AI gateway from LiteLLM to OmniRoute 2026-09-03 17:17:29 +00:00

Opened #32 OmniRoute: can it route to arbitrary OpenAI-compatible local endpoints (llama-server), not just Ollama? 2026-09-03 17:17:39 +00:00

Opened #33 OmniRoute due-diligence: maintainer, repo history, npm package trust 2026-09-03 17:17:47 +00:00

Opened #34 OmniRoute: memory/knowledgebase parity with litellm-pgvector 2026-09-03 17:17:58 +00:00

Opened #35 OmniRoute: web-search tool equivalent to the SearXNG standalone endpoint 2026-09-03 17:18:05 +00:00

Opened #36 OmniRoute: per-workload virtual keys, admin UI, and deployment shape (Docker/compose) 2026-09-03 17:18:10 +00:00

Opened #37 OmniRoute: docker-compose service design and secrets/config plan 2026-09-03 17:38:36 +00:00

Opened #38 Local image generation via ComfyUI + OmniRoute 2026-09-05 17:36:13 +00:00

Opened #39 Best diffusion model given full ~32GB headroom (Qwen stopped) 2026-09-05 17:36:34 +00:00

Opened #40 Why lazytainer idle-stop doesn't work today, and how the switch script should handle it 2026-09-05 17:36:35 +00:00

Opened #41 Add comfyui service to docker-compose.yml, register with OmniRoute 2026-09-05 17:36:37 +00:00

Opened #42 Extend downloader for the diffusion model weights 2026-09-05 17:36:38 +00:00

Opened #43 scripts/switch-model.sh: swap GPU residency between Qwen and ComfyUI 2026-09-05 17:36:39 +00:00

Opened #44 Add llama-server-fast: small non-thinking classifier/fast model for qwen-code Auto Mode 2026-09-06 18:17:27 +00:00