Open systems / 03
Open-source releases worth evaluating.
Find new AI tools and meaningful releases with source links, adoption context, and practical tradeoffs.
🎨 Open-source AI slide studio inside Codex: image-native decks, every slide a full visual canvas. ⚡ 10+ high-quality slides in ~4–5 minutes — Fast mode renders every page in parallel. 🔍 Watch the whole chain live: research → outline → style → render → edit → present → export P…
github.com2 months ago30 ptsView details
anthropics/anthropic-sdk-python v1.4.0
## 1.4.0 (2026-09-04) Full Changelog: [v1.3.0...v1.4.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.3.0...v1.4.0) ### Features * **api:** add Claude Tag category and user breakdowns to usage reports ([9fce1e4](https://github.com/anthropics/anthropic-sdk-python/…
github.com13 days agoView details
langchain-ai/langchain langchain-core==1.6.2
Changes since langchain-core==1.6.1 release(core): 1.6.2 (#40209) feat(openai): support async tools (#40208) chore(deps): bump mistune from 3.3.0 to 3.3.3 in /libs/core (#40150) chore(deps): bump tornado from 6.5.7 to 6.5.8 in /libs/core (#40113) fix(core): avoid mutation in goo…
github.com13 days agoView details
<details open> metal : add remaining fa-vec tunings for M3 (#28396) * addition of m3 in fa_vec_tuned_table * adding q4_0,q4_1,q5_0,q5_1 in ggml-metal-tuning * Fix formatting in ggml-metal-tuning.cpp </details> **Website:** - <https://llama.app> **Attestations:** - <https://githu…
github.com13 days agoView details
## Overview llama.cpp 0.4.0 adds initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support, on-demand tensor reading, per-slot server context limits, video input options, and a ggml update to 0.23.0 with major sparse flash attention and RDMA work. ### API changes - Added `llama_l…
github.com13 days agoView details
<details open> opencl: extend the elementwise and data‐movement op coverage (#27633) * opencl: add extended elementwise unary ops (sgn, step, elu, hardswish, hardsigmoid, floor, ceil, round, trunc) Adds nine GGML_UNARY_OP_* elementwise ops that were falling back to CPU on the Op…
github.com13 days agoView details
github.com2 months ago30 ptsView details
Deterministic-first execution engine for agent workflows in Go: the LLM extracts at the edge, a deterministic state machine decides. Zero dependencies, no-code JSON plugins, replayable runs.
github.com1 month ago20 ptsView details
Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.
github.com2 months ago27 ptsView details
https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
github.com3 months ago40 ptsView details
<details open> ggml-cpu(s390x) : fix q5_1 uninitialized v_acc (#28332) Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/45199911> **macOS/iOS:** - [macOS Apple Sili…
github.com14 days agoView details
<details open> src : add n_expert_used_max function (#28323) * src : add n_expert_used_max function With Commit c61b98b875eaa5e654a3f5c73b34c310d2c6ab4c ("model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444)") it is now possible for each layer to have a…
github.com14 days agoView details
<details open> sycl: fuse rms_norm+mul+add and add+add residual chains (#27610) Fuse RMS_NORM+MUL+ADD and ADD+ADD under GGML_SYCL_ENABLE_FUSION. ADD+ADD uses the same binbcast indexing and type matrix as standalone add() (f32, f16, f16/f32, i32, i16, bf16, including broadcast an…
github.com14 days agoView details
## [3.8.0](https://github.com/openai/openai-python/compare/v3.7.0...v3.8.0) (2026-09-03) ### Features * **api:** add gpt-6-astra and related features ([#3791](https://github.com/openai/openai-python/issues/3791)) ([09f446f](https://github.com/openai/openai-python/commit/09f446f5…
github.com14 days agoView details
面向 GPT-image-2 的 AI 图片生成 WebUI 工作台,支持 Codex Responses 与 OpenAI 兼容 API 接入,内置公用图库、多类型 Chip 快捷引用、提示词模板、多任务并发和本地队列管理。An AI image generation WebUI workbench for GPT-image-2 with Codex Responses and OpenAI-compatible API support, shared gallery references, multi-type quick chips, prom…
github.com3 months ago28 ptsView details
langchain-ai/langchain langchain==1.4.0
Changes since langchain==1.3.18 docs(langchain): runnable `langchain.mcp` examples (#39976) feat(langchain): `langchain.mcp` namespace, `MCPAdapter` (#39939) perf(anthropic,langchain): omit middleware trace inputs (#40098) fix(langchain): include model destination in agent tool…
github.com14 days agoView details
langchain-ai/langchain langchain-anthropic==1.7.1
Changes since langchain-anthropic==1.7.0 release(anthropic): 1.7.1 (#40181) perf(anthropic,langchain): omit middleware trace inputs (#40098) feat(anthropic): add Claude Fable 5.1 support (#40106)
github.com14 days agoView details
Free, local, open-source LLM cost analyzer - see where your LLM bill leaks, on your machine.
github.com2 months ago23 ptsView details
<details open> sycl: reduce redundant work in Q4_K multi-column MMVQ (#27062) * sycl: Q4_K Weight unpack optimization and reuse between destination Columns * sycl: Q4_K small N (N=2..4) + two output rows by subgroup reuse of activation between two rows. * sycl: gate Q4_K two-row…
github.com15 days agoView details
<details open> model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444) * hparams: add per-layer n_ff_exp/n_expert_used arrays with scalar-or-array loading G1/G2 infrastructure for variable-per-layer expert FFN size and top-k routing (required for Puzzle-75…
github.com15 days agoView details
QuesmaOrg/awesome-ai-tokenomics
A curated list of tools, benchmarks, papers, and copy-paste configs for AI token costs: what tokens cost, where they get wasted, and how to cut the bill.
github.com2 months ago22 ptsView details
<details open> mtmd: fix idefics3 preproc (#28273) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/44875193> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/downlo…
github.com15 days agoView details
The SQLite of agent sandboxes — self-hosted, E2B-compatible. One machine, sandboxes that live forever, idle costs nothing.
github.com2 months ago31 ptsView details
A 7-layer memory operating system for Hermes Agent — persistent memory with Qdrant, structured facts, fabric recall, auto-curated wiki, and surgical context injection. Runs locally, any LLM provider.
github.com3 months ago31 ptsView details
ZKE(Z Kubernetes Engine):AI 原生的 Kubernetes 云操作环境,桌面式多集群控制台加受控 AIOps Agent,基于 Server + Agent 与 QUIC/mTLS,适用于私有云、混合云及边缘环境
github.com1 month ago20 ptsView details
Local LLM cost-tracking proxy for OpenAI, Anthropic, Gemini, and pinned OpenRouter calls with token usage, failure, and billing-integrity receipts.
github.com2 months ago21 ptsView details
Self-hosted, Docker-first OpenAI-compatible LLM gateway with provider routing, fallbacks, API keys, and usage controls.
github.com2 months ago21 ptsView details
High-performance Knowledge Graph engine for AI, LLMs, and GraphRAG — built for the next generation of intelligent applications.
github.com2 months ago21 ptsView details
ADHD — a skill for coding agents. Tree-of-thought with pruning, built on the Claude & Codex Agent SDK. Fans out parallel divergent thoughts under different cognitive frames, scores, prunes traps, deepens the survivors. The no-brainer skill for creative and interdisciplinary work.
github.com3 months ago36 ptsView details
Local-first AI agent workspace for coding, writing, design, research, and automation — one runtime for desktop GUI and TUI.
github.com3 months ago38 ptsView details
Selfhost modern LLM stacks. Run the whole fleet from your terminal
github.com2 months ago20 ptsView details
<details open> vulkan : only request VK_KHR_shader_bfloat16 extension if supported (#28155) </details> **Website:** - <https://llama.app> **Attestations:** - <https://github.com/ggml-org/llama.cpp/attestations/44637403> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://gith…
github.com16 days agoView details
Self-hosted AI novel generator: single Go binary + web UI. OpenAI-compatible API → outline → chapter-by-chapter writing with review, foreshadowing, fact-check, and full-book polish. Chinese & English.
github.com3 months ago28 ptsView details
<details open> opencl: fix out‐of‐bound reads in the Adreno image kernels (#27632) * opencl: clamp the q4_K decode GEMV's fetch row on a padded x-grid * opencl: enforce the tiling contract of the image KQ/KQV GEMMs * opencl: decide the image KQ/KQV split at the dispatch, not fro…
github.com16 days agoView details
<details open> hexagon: add missing FARF logs for cpy/get_rows/set_rows/gdn ops (#28217) * hexagon: fix bug ne[2] printed in proc_op_req prep-src log * hexagon: add shape/VTCM farf logs to cpy, get/set rows, gdn </details> **Website:** - <https://llama.app> **Attestations:** - <…
github.com16 days agoView details
langchain-ai/langchain langchain==1.4.0a4
Initial release release(langchain): 1.4.0a4 test(langchain): cover mixed-era ClientGroup and group elicitation Update libs/langchain_v1/langchain/mcp/adapter.py fix(langchain): drive MCP elicitation via member session for fastmcp 4.0.1 fix(sdk): use latest fastmcp and rm reentra…
github.com16 days agoView details
World's fastest and most compact embedded vector database: exact by default, multimodal, local-first, and GPU-accelerated
github.com3 months ago20 ptsView details
A local multi-agent harness that works with your existing Claude Code, Codex subscriptions, allows you to run an office of agents
github.com3 months ago39 ptsView details
## [3.7.0](https://github.com/openai/openai-python/compare/v3.6.0...v3.7.0) (2026-09-02) ### Features * **api:** update usage APIs and documentation ([#3779](https://github.com/openai/openai-python/issues/3779)) ([6f0da16](https://github.com/openai/openai-python/commit/6f0da1657…
github.com16 days agoView details
Local proxy that compresses your LLM API requests so you pay less, with no change to the answers. Trims wasted tokens from prompts, history, tool output, and code before they're sent: -31% input / -74% output, measured live. Any provider, no extra model calls. Also an MCP server…
github.com3 months ago24 ptsView details
karthikreddy-7/ai-engineering-playbook
A zero-to-100 learning path for applied AI engineering — RAG, embeddings, vector search, agents, MCP, and the production engineering around them. 56 pages, built as a searchable site.
github.com1 month ago18 ptsView details
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
github.com3 months ago30 ptsView details
anthropics/anthropic-sdk-python v1.3.0
## 1.3.0 (2026-09-01) Full Changelog: [v1.2.0...v1.3.0](https://github.com/anthropics/anthropic-sdk-python/compare/v1.2.0...v1.3.0) ### Features * **api:** beta user profiles: add external_user_onboarded_at, remove relationship in favor of access_type ([74080c3](https://github.c…
github.com16 days agoView details
langchain-ai/langchain langchain==1.4.0a3
Third alpha of the `1.4.0` line. This release focuses on the new `langchain.mcp` namespace for adapting MCP servers into LangChain tools. ## `langchain.mcp` highlights - **`MCPAdapter`** adapts any target `fastmcp.Client` accepts — a URL, a local script, an in-process server, an…
github.com16 days agoView details
study8677/awesome-architecture
🧭 Architecture-first system design: 26 bilingual tutorials, 25 architecture templates, and 6 end-to-end cases covering distributed systems, AI-native systems, RAG, coding Agents, and production trade-offs.
github.com3 months ago34 ptsView details
fancyboi999/ai-engineering-from-scratch-zh
Agent工程师最全学习路径 · 从零精通 AI 工程 · 20 阶段 503 课 · 中文全量翻译 + 配套站点 + 动画讲解视频 · 如何成为 AI Agent 工程师的修成指南
github.com3 months ago30 ptsView details
如影随形 · 本地优先桌面工作日志:富文本记录、语义检索、工作台总结/问答;模型自配,数据留在本机。 Local-first desktop work journal—rich logs, semantic search, AI summary & Q&A. BYO models, data stays yours.
github.com3 months ago30 ptsView details
⚡ Awesome AI Gateway — pick an AI gateway from 160+ (LiteLLM, OpenRouter, Portkey, Kong, Higress, new-api, Bifrost) by cost, security & compliance, see what consolidated in 2026, learn how they actually work from an 8-chapter handbook, and verify every number yourself. Reproduci…
github.com3 months ago21 ptsView details
<details open> qwen4exp: support recurrent state rollback (#28123) MTP speculative decoding needs the target state to move back by the number of rejected draft tokens. Without rollback support the context is classified as SEQ_RM_TYPE_FULL and the server serializes the whole recu…
github.com17 days agoView details
Newsletter
Get practical AI engineering notes
Receive source-checked analysis of models, agents, evaluation, retrieval, and production reliability. Sent only when there is useful work to share.