Skip to content

Open systems / 03

Open-source releases worth evaluating.

Find new AI tools and meaningful releases with source links, adoption context, and practical tradeoffs.

  1. nexu-io/codex-slides

    🎨 Open-source AI slide studio inside Codex: image-native decks, every slide a full visual canvas. ⚡ 10+ high-quality slides in ~4–5 minutes — Fast mode renders every page in parallel. 🔍 Watch the whole chain live: research → outline → style → render → edit → present → export P…

    github.com2 months ago30 ptsView details

  2. anthropics/anthropic-sdk-python v1.4.0

    ## 1.4.0 (2026-09-04) Full Changelog: [v1.3.0...v1.4.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.3.0...v1.4.0) ### Features * **api:** add Claude Tag category and user breakdowns to usage reports ([9fce1e4](https://github.com/anthropics/anthropic-sdk-python/…

    github.com13 days agoView details

  3. langchain-ai/langchain langchain-core==1.6.2

    Changes since langchain-core==1.6.1 release(core): 1.6.2 (#40209) feat(openai): support async tools (#40208) chore(deps): bump mistune from 3.3.0 to 3.3.3 in /libs/core (#40150) chore(deps): bump tornado from 6.5.7 to 6.5.8 in /libs/core (#40113) fix(core): avoid mutation in goo…

    github.com13 days agoView details

  4. ggml-org/llama.cpp b10816

    <details open> metal : add remaining fa-vec tunings for M3 (#28396) * addition of m3 in fa_vec_tuned_table * adding q4_0,q4_1,q5_0,q5_1 in ggml-metal-tuning * Fix formatting in ggml-metal-tuning.cpp </details> **Website:** - <https://llama.app> **Attestations:** - <https://githu…

    github.com13 days agoView details

  5. ggml-org/llama.cpp v0.4.0

    ## Overview llama.cpp 0.4.0 adds initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support, on-demand tensor reading, per-slot server context limits, video input options, and a ggml update to 0.23.0 with major sparse flash attention and RDMA work. ### API changes - Added `llama_l…

    github.com13 days agoView details

  6. ggml-org/llama.cpp b10814

    <details open> opencl: extend the elementwise and data‐movement op coverage (#27633) * opencl: add extended elementwise unary ops (sgn, step, elu, hardswish, hardsigmoid, floor, ceil, round, trunc) Adds nine GGML_UNARY_OP_* elementwise ops that were falling back to CPU on the Op…

    github.com13 days agoView details

  7. mindscale-noah/MindMemOS

    github.com2 months ago30 ptsView details

  8. p0nymc1/cee

    Deterministic-first execution engine for agent workflows in Go: the LLM extracts at the edge, a deterministic state machine decides. Zero dependencies, no-code JSON plugins, replayable runs.

    github.com1 month ago20 ptsView details

  9. Orkas-AI/Orkas-VideoStudio

    Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.

    github.com2 months ago27 ptsView details

  10. StarTrail-org/PixelRAG

    https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/

    github.com3 months ago40 ptsView details

  11. ggml-org/llama.cpp b10797

    <details open> ggml-cpu(s390x) : fix q5_1 uninitialized v_acc (#28332) Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/45199911> **macOS/iOS:** - [macOS Apple Sili…

    github.com14 days agoView details

  12. ggml-org/llama.cpp b10796

    <details open> src : add n_expert_used_max function (#28323) * src : add n_expert_used_max function With Commit c61b98b875eaa5e654a3f5c73b34c310d2c6ab4c ("model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444)") it is now possible for each layer to have a…

    github.com14 days agoView details

  13. ggml-org/llama.cpp b10795

    <details open> sycl: fuse rms_norm+mul+add and add+add residual chains (#27610) Fuse RMS_NORM+MUL+ADD and ADD+ADD under GGML_SYCL_ENABLE_FUSION. ADD+ADD uses the same binbcast indexing and type matrix as standalone add() (f32, f16, f16/f32, i32, i16, bf16, including broadcast an…

    github.com14 days agoView details

  14. openai/openai-python v3.8.0

    ## [3.8.0](https://github.com/openai/openai-python/compare/v3.7.0...v3.8.0) (2026-09-03) ### Features * **api:** add gpt-6-astra and related features ([#3791](https://github.com/openai/openai-python/issues/3791)) ([09f446f](https://github.com/openai/openai-python/commit/09f446f5…

    github.com14 days agoView details

  15. kadevin/ilab-conjure

    面向 GPT-image-2 的 AI 图片生成 WebUI 工作台,支持 Codex Responses 与 OpenAI 兼容 API 接入,内置公用图库、多类型 Chip 快捷引用、提示词模板、多任务并发和本地队列管理。An AI image generation WebUI workbench for GPT-image-2 with Codex Responses and OpenAI-compatible API support, shared gallery references, multi-type quick chips, prom…

    github.com3 months ago28 ptsView details

  16. langchain-ai/langchain langchain==1.4.0

    Changes since langchain==1.3.18 docs(langchain): runnable `langchain.mcp` examples (#39976) feat(langchain): `langchain.mcp` namespace, `MCPAdapter` (#39939) perf(anthropic,langchain): omit middleware trace inputs (#40098) fix(langchain): include model destination in agent tool…

    github.com14 days agoView details

  17. langchain-ai/langchain langchain-anthropic==1.7.1

    Changes since langchain-anthropic==1.7.0 release(anthropic): 1.7.1 (#40181) perf(anthropic,langchain): omit middleware trace inputs (#40098) feat(anthropic): add Claude Fable 5.1 support (#40106)

    github.com14 days agoView details

  18. Rodiun/frugon

    Free, local, open-source LLM cost analyzer - see where your LLM bill leaks, on your machine.

    github.com2 months ago23 ptsView details

  19. ggml-org/llama.cpp b10777

    <details open> sycl: reduce redundant work in Q4_K multi-column MMVQ (#27062) * sycl: Q4_K Weight unpack optimization and reuse between destination Columns * sycl: Q4_K small N (N=2..4) + two output rows by subgroup reuse of activation between two rows. * sycl: gate Q4_K two-row…

    github.com15 days agoView details

  20. ggml-org/llama.cpp b10776

    <details open> model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) * hparams: add per-layer n_ff_exp/n_expert_used arrays with scalar-or-array loading G1/G2 infrastructure for variable-per-layer expert FFN size and top-k routing (required for Puzzle-75…

    github.com15 days agoView details

  21. tutti-os/tutti

    Where people and agents build in tune.

    github.com3 months ago36 ptsView details

  22. QuesmaOrg/awesome-ai-tokenomics

    A curated list of tools, benchmarks, papers, and copy-paste configs for AI token costs: what tokens cost, where they get wasted, and how to cut the bill.

    github.com2 months ago22 ptsView details

  23. ggml-org/llama.cpp b10775

    <details open> mtmd: fix idefics3 preproc (#28273) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/44875193> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/downlo…

    github.com15 days agoView details

  24. BitMiracle-AI/Dormice

    The SQLite of agent sandboxes — self-hosted, E2B-compatible. One machine, sandboxes that live forever, idle costs nothing.

    github.com2 months ago31 ptsView details

  25. ClaudioDrews/memory-os

    A 7-layer memory operating system for Hermes Agent — persistent memory with Qdrant, structured facts, fabric recall, auto-curated wiki, and surgical context injection. Runs locally, any LLM provider.

    github.com3 months ago31 ptsView details

  26. togettoyou/zke

    ZKE(Z Kubernetes Engine):AI 原生的 Kubernetes 云操作环境,桌面式多集群控制台加受控 AIOps Agent,基于 Server + Agent 与 QUIC/mTLS,适用于私有云、混合云及边缘环境

    github.com1 month ago20 ptsView details

  27. inferock/inferock-bench

    Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.

    github.com2 months ago21 ptsView details

  28. sina2266/Gozar

    Self-hosted, Docker-first OpenAI-compatible LLM gateway with provider routing, fallbacks, API keys, and usage controls.

    github.com2 months ago21 ptsView details

  29. Aliu-AiRobot/ESEILANE

    High-performance Knowledge Graph engine for AI, LLMs, and GraphRAG — built for the next generation of intelligent applications.

    github.com2 months ago21 ptsView details

  30. UditAkhourii/adhd

    ADHD — a skill for coding agents. Tree-of-thought with pruning, built on the Claude & Codex Agent SDK. Fans out parallel divergent thoughts under different cognitive frames, scores, prunes traps, deepens the survivors. The no-brainer skill for creative and interdisciplinary work.

    github.com3 months ago36 ptsView details

  31. KunAgent/Kun

    Local-first AI agent workspace for coding, writing, design, research, and automation — one runtime for desktop GUI and TUI.

    github.com3 months ago38 ptsView details

  32. raiyanyahya/llmaker

    Selfhost modern LLM stacks. Run the whole fleet from your terminal

    github.com2 months ago20 ptsView details

  33. ggml-org/llama.cpp b10756

    <details open> vulkan : only request VK_KHR_shader_bfloat16 extension if supported (#28155) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/44637403> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://gith…

    github.com16 days agoView details

  34. Nigh/show-me-the-story

    Self-hosted AI novel generator: single Go binary + web UI. OpenAI-compatible API → outline → chapter-by-chapter writing with review, foreshadowing, fact-check, and full-book polish. Chinese & English.

    github.com3 months ago28 ptsView details

  35. ggml-org/llama.cpp b10754

    <details open> opencl: fix out‐of‐bound reads in the Adreno image kernels (#27632) * opencl: clamp the q4_K decode GEMV's fetch row on a padded x-grid * opencl: enforce the tiling contract of the image KQ/KQV GEMMs * opencl: decide the image KQ/KQV split at the dispatch, not fro…

    github.com16 days agoView details

  36. ggml-org/llama.cpp b10753

    <details open> hexagon: add missing FARF logs for cpy/get_rows/set_rows/gdn ops (#28217) * hexagon: fix bug ne[2] printed in proc_op_req prep-src log * hexagon: add shape/VTCM farf logs to cpy, get/set rows, gdn </details> **Website:** - <https://llama.app> **Attestations:** - <…

    github.com16 days agoView details

  37. langchain-ai/langchain langchain==1.4.0a4

    Initial release release(langchain): 1.4.0a4 test(langchain): cover mixed-era ClientGroup and group elicitation Update libs/langchain_v1/langchain/mcp/adapter.py fix(langchain): drive MCP elicitation via member session for fastmcp 4.0.1 fix(sdk): use latest fastmcp and rm reentra…

    github.com16 days agoView details

  38. Egoist-Machines/LodeDB

    World's fastest and most compact embedded vector database: exact by default, multimodal, local-first, and GPU-accelerated

    github.com3 months ago20 ptsView details

  39. chaitanyagiri/munder-difflin

    A local multi-agent harness that works with your existing Claude Code, Codex subscriptions, allows you to run an office of agents

    github.com3 months ago39 ptsView details

  40. openai/openai-python v3.7.0

    ## [3.7.0](https://github.com/openai/openai-python/compare/v3.6.0...v3.7.0) (2026-09-02) ### Features * **api:** update usage APIs and documentation ([#3779](https://github.com/openai/openai-python/issues/3779)) ([6f0da16](https://github.com/openai/openai-python/commit/6f0da1657…

    github.com16 days agoView details

  41. fkiene/llmtrim

    Local proxy that compresses your LLM API requests so you pay less, with no change to the answers. Trims wasted tokens from prompts, history, tool output, and code before they're sent: -31% input / -74% output, measured live. Any provider, no extra model calls. Also an MCP server…

    github.com3 months ago24 ptsView details

  42. karthikreddy-7/ai-engineering-playbook

    A zero-to-100 learning path for applied AI engineering — RAG, embeddings, vector search, agents, MCP, and the production engineering around them. 56 pages, built as a searchable site.

    github.com1 month ago18 ptsView details

  43. Jwuthri/Tracely

    Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

    github.com3 months ago30 ptsView details

  44. anthropics/anthropic-sdk-python v1.3.0

    ## 1.3.0 (2026-09-01) Full Changelog: [v1.2.0...v1.3.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.2.0...v1.3.0) ### Features * **api:** beta user profiles: add external_user_onboarded_at, remove relationship in favor of access_type ([74080c3](https://github.c…

    github.com16 days agoView details

  45. langchain-ai/langchain langchain==1.4.0a3

    Third alpha of the `1.4.0` line. This release focuses on the new `langchain.mcp` namespace for adapting MCP servers into LangChain tools. ## `langchain.mcp` highlights - **`MCPAdapter`** adapts any target `fastmcp.Client` accepts — a URL, a local script, an in-process server, an…

    github.com16 days agoView details

  46. study8677/awesome-architecture

    🧭 Architecture-first system design: 26 bilingual tutorials, 25 architecture templates, and 6 end-to-end cases covering distributed systems, AI-native systems, RAG, coding Agents, and production trade-offs.

    github.com3 months ago34 ptsView details

  47. fancyboi999/ai-engineering-from-scratch-zh

    Agent工程师最全学习路径 · 从零精通 AI 工程 · 20 阶段 503 课 · 中文全量翻译 + 配套站点 + 动画讲解视频 · 如何成为 AI Agent 工程师的修成指南

    github.com3 months ago30 ptsView details

  48. FutureUniant/WorkShadow

    如影随形 · 本地优先桌面工作日志:富文本记录、语义检索、工作台总结/问答;模型自配,数据留在本机。 Local-first desktop work journal—rich logs, semantic search, AI summary & Q&A. BYO models, data stays yours.

    github.com3 months ago30 ptsView details

  49. cuihuan/awesome-ai-gateway

    ⚡ Awesome AI Gateway — pick an AI gateway from 160+ (LiteLLM, OpenRouter, Portkey, Kong, Higress, new-api, Bifrost) by cost, security & compliance, see what consolidated in 2026, learn how they actually work from an 8-chapter handbook, and verify every number yourself. Reproduci…

    github.com3 months ago21 ptsView details

  50. ggml-org/llama.cpp b10731

    <details open> qwen4exp: support recurrent state rollback (#28123) MTP speculative decoding needs the target state to move back by the number of rejected draft tokens. Without rollback support the context is classified as SEQ_RM_TYPE_FULL and the server serializes the whole recu…

    github.com17 days agoView details

Newsletter

Get practical AI engineering notes

Receive source-checked analysis of models, agents, evaluation, retrieval, and production reliability. Sent only when there is useful work to share.