vLLM: PagedAttention and Continuous Batching at Web Scale - Inside vllm-proje...
Most inference libraries treat GPU memory as one more thing for you to manage. vLLM, Apache-2.0-licensed and originally developed in the Sky Computing Lab at UC Berkeley, is the serving engine that made memory management the product: its PagedAttention scheme,...
SGLang: RadixAttention and Cache-Aware Serving - Inside sgl-project/sglang
Serving systems usually rediscover the same insight over and over: most tokens you generate have been generated before. SGLang, Apache-2.0-licensed and hosted by the LMSYS non-profit, turned that insight into its founding idea. The project grew out of the RadixAttention...
Qdrant: A Vector Database Written in Rust - Inside qdrant/qdrant
Embeddings are only half of a semantic search system; the other half is a database that can store billions of them, filter them by arbitrary metadata, and return nearest neighbors in milliseconds under concurrent writes. Qdrant, an Apache-2.0 vector database...
Open WebUI: A Self-Hosted Home for AI - Inside open-webui/open-webui
Run one container and you get a ChatGPT-shaped interface pointed at models you control: that is the promise that made Open WebUI one of the most starred projects on GitHub. The README calls it a home for AI — a...
Mem0: The Memory Layer That Gives Agents Recall - Inside mem0ai/mem0
Every LLM application rediscovers the same problem: the model knows everything about the world and nothing about your user. Context windows help for a turn or two, then the conversation moves on and yesterday’s preferences vanish. Mem0, an Apache-2.0 project...
LiteLLM: One Gateway for a Hundred LLM Providers - Inside BerriAI/litellm
Every team that adopts LLMs hits the same wall within months: a dozen provider SDKs, a dozen auth styles, a dozen error taxonomies, and no idea what any of it costs. LiteLLM, at version 1.106.0 and backed by Y Combinator,...
Letta Code: Agents That Rewrite Their Own Memory - Inside letta-ai/letta-code
Most coding agents wake up with amnesia every session: context is assembled, work is done, and everything learned evaporates when the process exits. Letta Code, from the team behind the MemGPT research paper, attacks that premise directly. Its README describes...
LangGraph: Pregel for Agents - Inside langchain-ai/langgraph
Agent frameworks usually promise simplicity and deliver fragility: the moment a workflow needs retries, approval gates, or state that survives a crash, the happy-path abstraction collapses. LangGraph, MIT-licensed and built by LangChain, starts from the opposite end. The README calls...
Haystack 3: Orchestration for Production RAG and Agents - Inside deepset-ai/h...
Frameworks for LLM applications tend to fail in one of two directions: too much magic, so nothing can be debugged, or too little structure, so every project reinvents plumbing. Haystack, the Apache-2.0 orchestration framework from deepset, has spent years threading...
Chroma: The Embedded Vector Database for AI Apps - Inside chroma-core/chroma
Some databases earn adoption by being effortless, and Chroma is the clearest example in the vector search world. Its README reduces the entire workflow to four functions — create a collection, add documents with metadata, query by text, and get...
VectorCraft: Exact Curve Booleans and 27 ms Renders - Inside storytold/vector...
Vector editors are where precision goes to die. Ask any tool to subtract one blob from another and half the time you get the infamous “cannot perform operation” dialog, a snapped anchor you never moved, or a file that renders...
Semantic Kernel: The Kernel That Ran Enterprise AI Agents - Inside microsoft/...
Before “agent framework” was a product category, Microsoft shipped an SDK with a stranger and better idea: treat the LLM like an operating system kernel. Prompts became functions, your Python methods became plugins, and a single orchestrator scheduled them all....
PhotoCraft: Photoshop Reborn in Pure Rust - Inside storytold/photocraft
Photoshop is the most-cloned application in creative software, and almost every clone has failed the same way: they rebuild the toolbar but not the object model, so layers stop behaving the moment you push past trivial edits. PhotoCraft, from the...
PdfCraft: An Acrobat Rebuilt From ISO 32000 - Inside storytold/pdfcraft
Acrobat is the app everyone uses and almost nobody trusts, because PDF tooling has a history of subscription walls, telemetry, and destructive saves. PdfCraft, part of the ArtCraft family on getartcraft.com, is the clean-room answer: an open-source reimplementation of Acrobat...
MetaGPT: A Software Company of LLM Agents in One Line - Inside geekan/MetaGPT
Most agent frameworks give you a loop and a toolbox. MetaGPT gives you an org chart. The repository at geekan/MetaGPT, MIT-licensed and currently at version 1.0.0, sells itself in one sentence: the multi-agent framework, and internally it includes product managers,...
LightCraft: A Lightroom Rebuilt in Pure Rust - Inside storytold/lightcraft
Lightroom has owned photo workflow for two decades, and every attempt to replace it has stumbled on the same wall: the library. Browsers can mimic the develop sliders, but ratings, flags, albums, previews, and the guarantee that your edits survive...
Langfuse: LLM Observability Built on ClickHouse - Inside langfuse/langfuse
When an LLM application misbehaves in production, the question is never whether something went wrong but where, and that question needs infrastructure to answer. Langfuse, the MIT-licensed open source LLM engineering platform at langfuse/langfuse, is that infrastructure: tracing, evaluation, prompt...
Jan: Rebuilding a Local ChatGPT Replacement on Tauri - Inside janhq/jan
Running an LLM on your own machine used to mean choosing between a CLI tool and a science project. Jan, self-described as an open-source ChatGPT replacement, occupies the middle ground: a real desktop product, downloadable from the Microsoft Store, Flathub,...
GraphRAG: When Your RAG Needs a Map - Inside microsoft/graphrag
Plain RAG retrieves passages; GraphRAG retrieves a map. The project at microsoft/graphrag, MIT-licensed and currently at version 3.3.0, describes itself as a data transformation suite that extracts meaningful, structured data from unstructured text using the power of LLMs, and the...
GPT Pilot: The Agent Team That Built Apps Step by Step - Inside Pythagora-io/...
Before the current wave of coding agents, there was GPT Pilot, the project at Pythagora-io/gpt-pilot whose thesis was refreshingly specific: AI can write most of the code for an app, maybe ninety-five percent, but a developer is needed for the...
FilmCraft: Premiere Pro Without FFmpeg - Inside storytold/filmcraft
Video editing is the hardest app to reimplement because it is really three apps stacked: an editing database, a media decoder zoo, and a rendering engine. FilmCraft, from the ArtCraft family on getartcraft.com, takes all three head-on in pure Rust....
EffectCraft: 306 Effects and Zero FFmpeg - Inside storytold/effectcraft
Every motion designer knows the deal: the compositor that defines the industry asks for a subscription, a login, and a machine that eats RAM for breakfast, and its project files are a binary blob you cannot even diff. EffectCraft, from...
DSPy: Programming, Not Prompting, Your Language Models - Inside stanfordnlp/dspy
Most prompt frameworks hand you a string and wish you luck. DSPy, from Stanford NLP, hands you a compiler mindset instead: you declare what goes in and what comes out with typed Signatures, compose those calls into Python Modules, and...
Knuth-Plass Type and Print-Grade PDF - Inside storytold/designcraft
Page layout is the most conservative craft in software. A magazine does not care how modern your toolkit is; it cares that the rag is even, the hyphenation is defensible, the folios land on the parent page, and the PDF...