Commit Graph

  • 2ea95fb059 fix cuda toolkit lookup and parallel (#16613) main Mark Ward 2026-06-30 12:56:54 -05:00
  • 8e7be3aed1 ci: avoid unbounded parallelism (#16966) Daniel Hiltgen 2026-06-30 10:49:55 -07:00
  • 710292ff4f mlx: tighten up gemma4 moe loading code (#16964) Patrick Devine 2026-06-29 21:15:08 -07:00
  • ada1eb5163 launch: check for min version for hermes desktop (#16912) Bruce MacDonald 2026-06-29 11:50:11 -07:00
  • 1c5ebbf5f4 llama.cpp update (#16960) Daniel Hiltgen 2026-06-29 09:43:41 -07:00
  • 7926b99e0e mlx: bump dependency (#16935) Daniel Hiltgen 2026-06-29 09:39:11 -07:00
  • 32a97b7493 tools: ignore braces inside JSON strings when detecting tool call end (#16937) Aditya Aggarwal 2026-06-28 00:30:55 +05:30
  • d26a58557d MLX: wire up scheduler selected context size for ps (#16918) Daniel Hiltgen 2026-06-26 08:47:03 -07:00
  • 2e474c98f9 parser/renderer: add Ornith 9B renderer/parser support (#16920) Parth Sareen 2026-06-25 23:18:47 -07:00
  • 2cb2c5381f launch: update hermes install urls to official (#16913) Bruce MacDonald 2026-06-25 16:22:19 -07:00
  • 2a6b50421a fix capability grid dark mode style (#16907) Eva H 2026-06-25 10:55:39 -07:00
  • f22ec2ec49 CUDA: require driver 550 or newer for v12 (#16895) Daniel Hiltgen 2026-06-25 08:46:00 -07:00
  • d9075caf1a docs: redesign coding integration docs (#16808) Eva H 2026-06-25 07:03:59 -07:00
  • e11eeb3ba0 llama.cpp version update (#16548) Daniel Hiltgen 2026-06-24 14:03:12 -07:00
  • 0a408b2225 jetson: add CC 87 for CUDA v13 (#16628) Daniel Hiltgen 2026-06-24 14:02:41 -07:00
  • 16739dee60 server: align generate with native chat templates (#16878) Daniel Hiltgen 2026-06-24 13:43:56 -07:00
  • d48d790baf docs: redesign docs landing and integrations overview (#16807) Eva H 2026-06-24 13:28:28 -07:00
  • 0463940334 llm: fix ollama ps double-counting mmap'd weights on partial offload (#16709) Philip Sinitsin 2026-06-24 19:43:20 +01:00
  • 570679c9e0 mlx: update and fix CUDA JIT packaging (#16871) Daniel Hiltgen 2026-06-24 10:36:02 -07:00
  • 89a171cc70 llm: use host Vulkan loader on Windows (#16869) Daniel Hiltgen 2026-06-24 10:35:48 -07:00
  • 33878e671a llama: default qwen2.5vl window attention metadata (#16868) Daniel Hiltgen 2026-06-24 10:35:29 -07:00
  • c191a145bb llm: preserve generation headroom for shifted prompts (#16856) Parth Sareen 2026-06-23 15:29:40 -07:00
  • 479e1cf94e docs: document max think level (#16877) Parth Sareen 2026-06-23 15:29:15 -07:00
  • 836507378b llm: size mmproj offload by projector memory (#16866) Daniel Hiltgen 2026-06-23 13:04:02 -07:00
  • 46bc1bcb4c llama: add sm_86 architecture to cuda_v13_windows preset (#16834) anish 2026-06-23 07:35:21 -07:00
  • 2a8b31531e launch/codex: detect model drift when Codex App UI switches away from Ollama (#16864) Bruce MacDonald 2026-06-22 15:38:19 -07:00
  • 505e35f2b9 mlxrunner: choose the speculative draft length to maximize throughput Jesse Gross 2026-06-09 17:17:26 -07:00
  • 114875133b mlxrunner: resolve each speculative round in one host sync Jesse Gross 2026-06-11 14:45:35 -07:00
  • 42c330283b mlxrunner: run one target forward per MTP decode step Jesse Gross 2026-06-09 17:14:30 -07:00
  • f93efe2809 mlxrunner: apply in-flight drafts to proposal penalty history Jesse Gross 2026-06-12 16:17:21 -07:00
  • 28fbbb06d5 mlxrunner: support draft heads that maintain draft caches Jesse Gross 2026-06-09 17:10:40 -07:00
  • 340c51bbb7 mlxrunner: host speculative decoding in the text generation pipeline Jesse Gross 2026-06-09 16:56:59 -07:00
  • 2e9d68dc38 mlxrunner: unify the MTP decode paths Jesse Gross 2026-06-02 12:53:47 -07:00
  • fc58544422 discover: fix inverted iGPU/dGPU Vulkan classification on Windows hybrid graphics (#16669) Sahil Kadadekar 2026-06-22 17:52:03 -04:00
  • e434a93884 launch: auto-install opencode when missing (#16806) Eva H 2026-06-19 10:12:11 -07:00
  • 9c02d8e69d launch: auto-install Claude Code (#16802) Eva H 2026-06-19 10:11:50 -07:00
  • 07ed752353 launch: add thinking capability detection to opencode (#15434) Eva H 2026-06-18 10:45:16 -07:00
  • e1f7f9cbdb ci: pin darwin release xcode (#16788) Parth Sareen 2026-06-17 13:01:10 -07:00
  • 8c432fc88a llama: update llama.cpp to b9672 (#16775) Patrick Devine 2026-06-16 23:15:52 -07:00
  • acfb50d9af models: add cohere2_moe (Command A / North) to the MLX engine (#16670) Jeffrey Morgan 2026-06-16 23:15:21 -07:00
  • 0f047feef5 llm: context shift allow shiftable prompts (#16764) Jeffrey Morgan 2026-06-16 12:55:52 -07:00
  • 9e4ed74efe integration: look for the "hf" tool in integration tests (#16765) Patrick Devine 2026-06-16 11:04:54 -07:00
  • bbb40a0a6c server: context shift for context windows larger than 8k, add error when hitting context limit (#16712) Jeffrey Morgan 2026-06-15 11:36:50 -07:00
  • 993acc7504 model: update lfm2 parser/renderer for optional thinking (#16359) Jeffrey Morgan 2026-06-14 20:37:08 -07:00
  • 7ea692cb2b llama: update llama.cpp to b9637 (#16609) Jeffrey Morgan 2026-06-14 20:05:08 -07:00
  • 12e04379cd launch: Fix launch provider drift (#16683) Parth Sareen 2026-06-11 17:21:46 -07:00
  • f8a48df24d llm: decouple prompt caching from context shift (#16639) Parafee41 2026-06-12 07:05:24 +08:00
  • 82e0ddb6fe mlxrunner: harden linear/embedding layers against over-promotion (#16682) Patrick Devine 2026-06-11 13:56:25 -07:00
  • 1abd56b6e6 mlxrunner: record committed MTP drafts before streaming them Jesse Gross 2026-06-01 11:00:46 -07:00
  • ded2db7d86 mlxrunner: capture prefill snapshots across the forward Jesse Gross 2026-05-29 12:12:10 -07:00
  • d00622060f mlxrunner: drive MTP speculation through cache snapshots Jesse Gross 2026-05-29 12:11:30 -07:00
  • 177aefb8a9 nn/recurrent: return per-boundary states from the gated-delta kernels Jesse Gross 2026-05-29 12:07:46 -07:00
  • 07588c64ee mlxrunner/cache: split KVCache and RotatingKVCache into their own files Jesse Gross 2026-05-29 13:16:27 -07:00
  • 4c97a940ca mlxthread: preserve the original stack when worker work panics Jesse Gross 2026-05-27 13:07:33 -07:00
  • 74cbf1d2c2 docs: omp (#16552) Bruce MacDonald 2026-06-08 11:43:51 -07:00
  • 5c1e37eb67 docs: hermes desktop (#16549) Bruce MacDonald 2026-06-08 11:43:11 -07:00
  • f0078ae476 docs: update docs examples to use Gemma 4 instead of Gemma 3 (#16607) Jeffrey Morgan 2026-06-07 12:43:13 -07:00
  • 96201a623a Add AGENTS.md and CLAUDE.md to root repository (#16604) Jeffrey Morgan 2026-06-07 10:57:59 -07:00
  • 9c94c2b11e docs: describe llama.cpp update process (#16603) Daniel Hiltgen 2026-06-07 10:27:47 -07:00
  • e09b3f9fb5 openai: align models list with tags (#16556) Parth Sareen 2026-06-05 17:59:05 -07:00
  • a0099da2d1 launch: use native Windows Hermes config path (#16558) Bruce MacDonald 2026-06-05 17:29:19 -07:00
  • 25e0e81e12 docs: update Zod example to use native toJSONSchema (#14746) Chris Chen 2026-06-06 09:21:07 +10:00
  • 87cff95af8 launch: oh-my-pi (#16410) Bruce MacDonald 2026-06-04 17:49:49 -07:00
  • 3ef69ef784 mlx: allow the embedding layer to use the nvfp4 global scale (#16527) Patrick Devine 2026-06-04 17:40:01 -07:00
  • 1a7786be14 docs: add cloud model retirement (#16528) Michael Yang 2026-06-04 15:18:38 -07:00
  • 3370ff8b1c launch: hermes-desktop app (#16516) Bruce MacDonald 2026-06-04 11:51:36 -07:00
  • 455f57457d llama.cpp version update (#16511) Daniel Hiltgen 2026-06-04 08:20:57 -07:00
  • 1d955ed990 integrations: hermes windows install (#16487) Bruce MacDonald 2026-06-03 17:40:45 -07:00
  • d071237131 docs: add Cline CLI integration doc (#16341) Eva H 2026-06-03 20:30:01 -04:00
  • 229a1303fb llama-server: fix gemma4 patch wiring (#16477) Daniel Hiltgen 2026-06-03 14:41:03 -07:00
  • ac3d0657a2 launch: migrate pi (#16213) Parth Sareen 2026-06-03 14:35:32 -07:00
  • 01557ff313 llama-server: allow GPU offload for projectors (#16473) Daniel Hiltgen 2026-06-03 13:58:40 -07:00
  • e5a38739b4 mlx: "requires" in modelfile is being ignored for mlx based models (#16469) Patrick Devine 2026-06-03 13:10:57 -07:00
  • 5f56a289b3 server: classify mmproj GGUFs as projector layers (#16472) Jeffrey Morgan 2026-06-03 12:59:34 -07:00
  • ad8cda255d launch: clean legacy codex profile before launch (#16467) Eva H 2026-06-03 14:49:31 -04:00
  • 3e1b4fe39d Kill llama-server during Windows cleanup (#16458) Daniel Hiltgen 2026-06-03 10:25:12 -07:00
  • 52196f1a97 llama.cpp version update (#16463) Daniel Hiltgen 2026-06-03 10:20:30 -07:00
  • 50bbda5660 models: add support for gemma4-12b (#16457) Patrick Devine 2026-06-03 07:44:57 -07:00
  • 4b5bdd3b25 fix laguna patch build breakage (#16445) Daniel Hiltgen 2026-06-02 16:35:19 -07:00
  • e828061b6e llm: ignore llama-server SSE ping comments (#16443) Daniel Hiltgen 2026-06-02 15:40:14 -07:00
  • 7a2073d17b docs: configure hermes desktop app (#16440) Bruce MacDonald 2026-06-02 14:32:10 -07:00
  • c952708169 llama: add laguna (poolside) arch via a llama.cpp patch under llama/c… (#16396) Daniel Hiltgen 2026-06-02 13:17:08 -07:00
  • f57d111754 launch: isolate Codex launch configuration (#16437) Parth Sareen 2026-06-02 12:10:46 -07:00
  • c34a79a373 llama.cpp version update (#16426) Daniel Hiltgen 2026-06-02 11:46:56 -07:00
  • b051c9cf83 More harden app markdown URL handling (#16436) Daniel Hiltgen 2026-06-02 11:46:14 -07:00
  • b7b7fa0454 llm: detect llama-server load stalls from output (#16427) Daniel Hiltgen 2026-06-02 11:30:48 -07:00
  • 4c076813be discover: allow Radeon 8060S iGPU by default (#16429) Daniel Hiltgen 2026-06-02 11:15:01 -07:00
  • 6780f0416a Harden app markdown URL handling (#16380) Daniel Hiltgen 2026-06-02 11:14:36 -07:00
  • 35fa277fa9 llm: include cached prompt tokens in llama-server counts (#16428) Daniel Hiltgen 2026-06-02 10:51:01 -07:00
  • 05747b02ab launch: fix opencode local model limits (#16425) Daniel Hiltgen 2026-06-02 10:50:35 -07:00
  • 4e807fdedd cmd/launch: add Qwen code integration (#15900) Eva H 2026-06-01 19:51:32 -04:00
  • 7d3a6c3ae5 log template details to aid troubleshooting (#16403) Daniel Hiltgen 2026-06-01 16:25:44 -07:00
  • 06ff728246 feat(launch): show and auto-install Cline CLI (#16402) Eva H 2026-06-01 19:04:08 -04:00
  • 2c71d8d7ca launch: migrate Codex config (#16397) Parth Sareen 2026-06-01 13:46:41 -07:00
  • 00381496a3 cmd/launch: fix configure cline ollama provider via providers.json (#16352) Eva H 2026-06-01 16:41:40 -04:00
  • 5e9636fa05 launch: avoid legacy Codex App profiles (#16364) ZiTian Zhao 2026-06-02 02:49:56 +08:00
  • 630882621b llama-server followups (#16353) Daniel Hiltgen 2026-06-01 10:44:21 -07:00
  • 0e93ccc2cd convert: fixes for qwen3next model conversion (#16354) Patrick Devine 2026-06-01 09:43:11 -07:00
  • e7766a4a47 model: improvements to laguna-xs.2 parser/renderer (#16362) Jeffrey Morgan 2026-05-31 14:11:07 -07:00
  • be7de10c41 llama: handle Gemma 4 and LFM2 BOS override in llama server (#16367) Jeffrey Morgan 2026-05-31 14:05:39 -07:00