Skip to content
toolingsystems

After Git · Part 1 of 3

Git Won. Version Control Isn’t Finished.

loadingalias

· post · 20 min

Git won on speed, distribution, and compatibility. The next VCS must version whole projects, preserve intent, and migrate from Git instead of wrapping it.

Git doesn’t suck. That would be an easier article to write, and a much less useful one.

Git solved its problems so well that it became the compatibility layer for software development. Editors understand it. Build systems understand it. Every serious forge understands it. Entire companies exist to host it, accelerate it, secure it, and hide its sharp edges.

That success has made Git’s limitations easier to live with. However, I’m still of the opinion that Git needs to be replaced, entirely, and that means taking on an enormous engineering job.

Git can store almost any project file, and partial clones can fetch missing objects on demand. What it lacks is a native model for assets with different transfer, editing, and retention policies; changes that survive rewrites; several agents attempting the same task; and the evidence that explains why one attempt failed. We need a system of record not just for the code, but for the path taken to get there.

As developers, we hack around it; we encode that state in branch names, PRs, external storage, task databases, chat logs, worktrees, and our own brains. We’ve seen a handful of tools pop up to work around these limitations for agentic workflows. Git remains underneath all of it. Compatibility (and fear of even trying) keeps winning, even when the resulting workflow is a mess.

A month or so ago, a reply I posted on X about the future of Git introduced me to a brilliant mind whose ideas felt similar to my own. He pushed me to name the actual bottleneck, then pointed out how Fossil’s colocated project history could help agents. I hadn’t considered that use for it. This led me to reconsider the work I’d done in this space five years earlier. It reignited my passion for a better version control system.

I want to explain why Git won, where its model stops, and which primitives a successor actually needs. Replacing the forge - GitHub, GitLab, CodeBerg, etc. - is a separate problem, which I’ll cover in Part Two. Make no mistake, there are dragons along both paths.

Git stores states exceptionally well

Git becomes a lot easier to wrap your head around if you ignore most of its command surface.

At its core, it’s just a content-addressed object database:

  • A blob contains file content.
  • A tree maps names to blobs and other trees, recursively representing a repo’s state.
  • A commit references a tree and parent commits, and contains author, committer, and message fields.
  • A ref is a mutable name pointing to an object. Branches and tags are built from refs.

A branch doesn’t contain commits. It points into a commit graph.

In this hypothetical one-file repo, README.md changes from hello in C0 to hello, world in C1:

branch main
    │
    ▼
commit C1 ── parent ──▶ commit C0
    │                       │
    ▼                       ▼
  tree T1                 tree T0
    │                       │
    ▼                       ▼
README.md → blob B1     README.md → blob B0
"hello, world"          "hello"

change Δ = diff(T0, T1)

Git stores the commits, trees, and blobs. It calculates Δ by comparing T0 and T1; the change isn’t a durable logical object that survives when its implementation is replaced. Git’s data-model documentation describes the objects in much more detail than I do here.

Git also maintains an index, which we know (and universally loathe) as the staging area. Internally, it’s a flat list of paths and object IDs used to construct a tree and represent merge stages. This is important; it’s necessary, even. The issue is that the boundary escaped into everyday use, leaving devs to reason about a second mutable state between the working tree and history.

On disk and over the network, Git groups objects into packfiles. Blobs are logically complete values, but packed objects may be compressed as deltas against other objects. Fetch and push negotiate missing objects and transfer packs. The format is compact, very mature, and highly optimized.

This is a genuinely formidable foundation. There is a reason, outside of fighting adoption, that Git is the universal truth for all of our work… and it starts here. Objects are immutable and verifiable. History is local. Work continues offline. Exact repo states are dirt cheap to name and exchange. Git has also had more than twenty years of perf and compat work by some of the smartest engineers on the planet.

Perhaps ironically, the limitations follow from that design. Git stores and exchanges repo states; we reconstruct intent, policy, attempts, and most of the project’s context around them.

That creates two different problems. Some project material needs storage policy Git doesn’t own. The work that moves one state to another needs an identity Git doesn’t preserve. Let’s start with the raw material.

Today’s VCS market is already split

The VCS market looks a lot simpler from a source repo than from a complete project. Git owns source-code dev. Asset-heavy teams still reach for P4 or Unity Version Control; projects with datasets or machine learning models add another layer. No current VCS is the obvious home for source and every durable input that produces it.

Git LFS replaces a large file in Git with a text pointer while storing the payload on another service. DVC keeps dataset and model placeholders in Git while DVC remotes hold the data in S3, Azure, GCS, SSH storage, etc. Hugging Face’s Xet-backed repos add chunk-level dedup and smaller transfers for large ML/DL files while preserving Git repositories as the interface.

As an aside, HuggingFace has solved or innovated around a handful of real issues Git’s presented the ML community. The same could be said of DeepSeek, as well. The 3FS system is brilliant. Anyone working on the future of VCS should definitely include the work of HuggingFace on Xet, and DeepSeek on 3FS.

With all three (Git LFS, DVC/Xet remotes), source history lives in one graph and large content lives elsewhere. Availability, auth, retention, garbage collection, transfer, and provenance cross systems. Metadata can survive after its payload disappears. A checkout may require Git, an extension, another remote, and credentials Git can’t describe.

We have normalized these splits because every alternative sacrifices part of Git’s ecosystem or its local-first (offline) behavior.

Those pointers already name exact versions. LFS records a payload hash and size; DVC records the identity of tracked data. The missing contract is what the VCS does with those references.

Any successful successor must make them native: publishing a revision verifies that required content is available under its durability policy, and retention and garbage collection follow references across storage paths. One logical history doesn’t keep payloads alive by itself. It needs lifecycle rules with teeth. A clone should distinguish content it has chosen not to fetch from content that has been lost.

That shared identity doesn’t mean storing every byte the same way. Text, a Blender scene, and a 40 GB video need different policies for chunking, merging, locking, and retention, as well as which bytes a workspace fetches and when. The latter of which is crucial to get right for a smooth DX. Uniform identity is useful; uniform encoding would be a huge mistake here.

Agents expose the missing unit of work

Even an ordinary source repo has the second problem: identifying the work that changes it.

Git already supports parallel work. Branches are cheap. Worktrees provide several working directories with separate HEAD and index state while sharing repo data. They isolate files. They don’t model the work, though.

Consider this simplified, hypothetical graph for one engineer supervising a few agents:

change A: refactor parser
├─ revision A1: minimal refactor [accepted]
└─ revision A2: larger rewrite [rejected]
   └─ observation O1
      └─ parser bench regressed

change B: improve diagnostics
├─ logically depends on change A
└─ revision B1 was built and tested against A1

change C: update docs
├─ no code dependency on A
└─ must describe accepted revision A1

An orchestrator agent can assign a branch and worktree to each attempt, stack B on A1, rerun checks, and rebase everything when main moves. That doesn’t tell it:

  • whether A1 and A2 are revisions of one change or unrelated branches;
  • whether B depends on change A or the current commit implementing it;
  • which observation rejected A2 and which dependent work it invalidated;
  • who owns the change when humans, agents, and sessions all acted on it.

Git can encode pieces of this with commits, notes, refs, and conventions. Gerrit’s Change-Id already connects revisions across amendments and rebases. Jujutsu (JJ) puts stable changes into the local VCS model. The requirement is that retained work history remain interpretable without the original forge, and that native operations preserve its relationships without every integration reconstructing them.

A branch called parser-refactor works great while one person knows what it means. When several engineers and their agents work on it concurrently, the name alone tells them essentially nothing; this is a nightmare today.

The missing unit is a durable change graph: stable work identities, immutable revisions, dependencies between changes, ownership, alternate attempts, observations, and the operation history that produced the accepted version.

An abandoned attempt may tell the next agent which constraint was violated or which bench regressed. Git preserves the commits we choose to keep. Agentic dev also benefits from preserving why we rejected the others. We have reached the point where that extra data, the story of how we got to this point, is as valuable as the current work is.

The design space never closed

VCS and adjacent tools before and after Git optimized for different things. The table below compares the primitives they made explicit and the problems they left open. We all know who won… but that doesn’t tell us what else was possible.

PrimitiveSystemsWhat it establishedBoundary that remained
Atomic project-wide revision and central authoritySubversionRepo revisions could be atomic, with efficient binary deltas, cheap copies, and optional lockingOne clear server authority makes the network part of ordinary work
Distributed local historyMercurialGit wasn’t the only credible model for complete local history and distributed workTechnical credibility doesn’t overcome an ecosystem that standardized elsewhere
Typed assets and edit policyPerforce P4 and Unity Version ControlLarge binaries, artist workflows, file-type policy, and exclusive checkout can belong inside version controlSpecialization preserves capabilities that Git-centric projects often move into another system
External data versioningDVCDataset and model snapshots can be tracked beside source without placing their payloads in GitGit metadata, local caches, and data remotes remain separate systems with separate lifecycles
Data-native version controlDoltTables and schemas can support commit, diff, branch, and merge over content-addressed storage designed for seek and structural sharingA specialized relational database isn’t a whole-project VCS
Durable changes and recoverable mutationJujutsu (changes and working copy, undo and operation log)Automatic snapshots, stable change IDs, committed conflicts, whole-repository undo, and safer history editing can materially improve the user modelIts Git backend is practical, but richer state doesn’t survive every Git operation
Patch dependency and representable conflictDarcs and PijulPatches or graph-native changes can carry dependencies, be applied in either order when independent, and preserve conflicts as dataPatch theory still needs a product and migration story strong enough to cross the ecosystem
Colocated project memoryFossilCode, tickets, wiki, forum, and technical notes can synchronize as one project historyAn all-in-one project model is a large adoption demand
Native and compatibility modesSaplingA system can separate Git-compatible operation from a stronger native directionSome native scale capabilities are unavailable in Git mode or not included in public releases

This is hardly exhaustive. There are glaring omissions. Nevertheless, the point is made.

These aren’t failed drafts of Git. P4 and Unity Version Control remain current because game and media projects still have requirements that source-oriented workflows handle poorly. DVC exposes the cost of keeping data identity in Git while the payload lifecycle lives elsewhere. Dolt shows that a data modality can justify a storage engine built for versioning rather than adapted to it. Jujutsu, Darcs, Pijul, Fossil, and Sapling make operations, changes, conflicts, project memory, or compatibility boundaries explicit instead of reconstructing all of them around commits.

None of these systems offers a complete route through semantics, assets, scale, migration, and ecosystem compatibility at the same time.

Git itself was built for the Linux kernel after the kernel project’s relationship with BitKeeper broke down in 2005. Its design goals included speed, simple design, full distribution, efficient handling of a kernel-sized project, and support for thousands of parallel branches.

Git won because it was fast, distributed, scriptable, open, and good enough for an enormous range of software projects. Then the advantage compounded. GitHub made Git just social enough to stick. IDEs, CI providers, package registries, deployment platforms, and review tools standardized around its refs and object IDs. Choosing Git stopped feeling like a choice. This brings us to the present.

That ecosystem is now Git’s greatest strength, and the moat every successor has to reckon with. In The Disappearing Moat, I argued that accumulated systems (standards, integrations, ecosystem position) are among the few advantages that survive cheap imitation. Git’s ecosystem is the textbook case. Git didn’t exhaust the design space; its ecosystem made looking beyond it harder to justify - nearly impossible.

Scaling Git isn’t the future

Cursor’s Continuity design is pretty serious systems work aimed at Git’s storage limits. It puts a WAL in S3-shaped storage at the authority boundary and treats ordinary Git repos on local NVMe as materialized caches. It sparked the “s3 shaped everything” moment we’ve seen lately, too. Continuity allows Cursor Origin to serve standard Git ops while handling huge monorepos and vast numbers of small agent repositories.

Cloudflare Artifacts makes a similar semantic choice for agent automation: durable, programmable repos behind a Git-compatible interface, reachable from Workers, a REST API, or standard Git clients.

Both projects solve operational problems while preserving Git’s interface. Continuity makes local repos disposable caches. But better replication doesn’t give a change an identity across rewrites or make asset references share one lifecycle contract. Git’s native semantics stay where they were.

A successor will still need replication, caching, compaction, routing, and repair; missing work identity didn’t create those distributed-systems costs. The question is which costs the workload requires and which come from translating it into a model that doesn’t represent it directly. I want the next VCS to remove work at the foundation: move fewer bytes, keep fewer redundant representations, and model the project directly.

East River Source Control goes further and pitches itself as what comes after Git. Its plan is to “speak the Git protocol, while changing how the storage layer works”: Git clients talk to a bridge, and behind it sits a custom storage engine with Jujutsu’s model, first-class conflicts, and fine-grained ACLs. jj’s creator, Martin von Zweigbergk, is now its CTO. At Google, he worked on Fig, the Mercurial client on top of Piper, and jj has a Piper backend there. East River is bringing that playbook to Git. jj and any client changes stay open source; the server side is proprietary.

I can appreciate the talent and desire for maximum interop with Git. I still think it’s the wrong shape. Serving the Git protocol keeps Git’s object model and identities at the center of the design: whatever the native model adds has to survive translation for unchanged Git clients or stay invisible to them. That’s a compatibility tax built into the foundation. It may be a better Git host. I don’t think it’s the next VCS.

Piper’s playbook also makes a central server the authority for history. I’m not convinced that’s the right direction either; local history and offline work are a big part of why Git won.

So no, I don’t think scaling Git, or rebuilding storage behind its protocol, is the future of version control. Continuity and Artifacts are useful compatibility bridges. Building a successor with compat handcuffs is the fastest way I know to fail to realize a zero-to-one success. I’d rather build the best VCS we can and migrate teams off Git very well, with a dedicated migration engine that lets us leave behind the legacy code we no longer need. This is not a trivial task; it’s also not the end of the world.

Four requirements for the next VCS

Five years ago, after thinking about this problem for years, I set out to replace GitHub. This was a real goal. That meant rethinking the VCS underneath it, and the work expanded across storage engines, content-defined chunking (CDC), sync protocols, patch theory, CRDTs, merge semantics, op logs, lazy materialization, and the boundary between a VCS and its forge. Some experiments were dead ends. Others changed what I thought the system had to own and how it should work.

I don’t have a version control system to publish, and I don’t know if I ever will. The work kept pulling me further down the stack. These requirements came out of it. They will fight each other, and storage alone remains an enormous problem. A successor still has to meet all four: project material, work identity, operability, and adoption.

One history, different data paths

I’ve settled on Git’s treatment of source files being mostly right. Code should stay whole logical blobs: mandatory app-level chunking adds indexes and garbage collection that small text files never repay. That’s a logical rule, not a physical one. Git already delta-compresses blobs inside packfiles while the logical object stays simple.

Datasets, model weights, game assets, video, and large generated inputs are different, which is exactly why DVC, P4, and Dolt exist.

Every durable input still needs first-class identity in one project history. A revision should name the exact code, datasets, models, assets, specs, and generated inputs that produced a result. A large mutable dataset will almost certainly benefit from modality-aware tuned CDC, record-aware partitions, immutable shards, or structural trees. Byte boundaries and logical partitions solve different problems. A 3D scene may require an exclusive lease or app-aware merge. Already-compressed, immutable media may still be best stored whole… but I think it needs to be explored deeply to make that call.

No algorithm makes every large object cheap. Dense model weights or compressed media may share too little reusable content for chunking to pay for its indexes, verification, and garbage collection. The physical policy has to follow the workload.

For large mutable data, I want the host to control the physical storage policy it accepts without defining logical file identity. Clients must execute enough of that policy locally to avoid uploading unchanged content. Duplicate data has to stay off the wire. A versioned protocol lets the host advertise its algorithm and parameters, the client derive chunk identities, and both sides reconcile before transferring payload. Offline revisions must remain possible, and two remotes may choose different representations of the same logical file. Rechunking between those policies has a cost the design must account for.

I’ve spent a long time here. I’ve explored combining CDC with an invertible Bloom lookup table (IBLT) handshake over QUIC. Chunking found content-derived byte boundaries. The IBLT reconciled the resulting sets of chunk identities. QUIC carried the missing content. It was promising, but sizing and retrying the reconciliation, verifying content, reclaiming chunks, and materializing partial files all become core VCS work. I don’t know if that’s the right design. CDC also needs tuning for the data; it won’t suit every workload. Even the best algos today will benefit from modality-aware tuning at scale.

I came away convinced that efficient transfer needs an explicit protocol for identity, reconciliation, verification, and reconstruction. Xet already proves that such a protocol can coexist with Git. My argument for native assets goes further: publication, retention, and recovery must understand those references without each tool rebuilding the contract around pointer files.

A forge built on one history for code, datasets, models, and assets would reach far beyond GitHub’s current market. I want it to compete for the workflows now split across Hugging Face and Kaggle. I want less fragmentation. I want fewer scattered pieces and in this case, more silos. Obviously, winning them still requires execution environments, discovery, collaboration, and distribution; storage architecture isn’t going to deliver those, nor will it cut it. Part Two will cover the production system that would have to exist to support them.

Changes and revisions must be separate

A change is durable intent. A revision is one immutable materialization of that change against an exact state.

Rebasing, refining, or asking another agent for a different implementation should create a revision without erasing the identity of the work. Logical deps connect changes; each materialized revision binds them to exact revisions. B can follow the continuing work called A while recording that B1 was tested against A1. When A changes, the relationship survives, but the test result doesn’t automatically apply to the new combo. Review, tests, provenance, and discussion can then attach to the right layer.

I spent a significant part of the work experimenting with patch theory and CRDTs. For a long time, I was convinced that patch theory was the only way forward. Patch models force deps, commutation, inversion, and conflict into the data model. Foundational CRDT research gives precise conditions for replicas to converge. Neither decides whether two edits to a schema or 3D scene are semantically compatible, though.

Those experiments have nearly convinced me that no single theory will settle the design. Concurrent work, alternate revisions, and conflicts should remain inspectable states instead of temporary failures that the repo can’t represent.

The VCS should also carry an append-only dev record linked to changes and revisions: actors, attempts, observations, decisions, and outcomes. The forge can attach collaboration and policy, but the retained dev record shouldn’t vanish when the remote does. This is really important for a variety of reasons I won’t bore you with, but the ability to work offline cannot be ignored. This needs to be a part of the core thesis.

Append-only describes mutation semantics while a record is retained, not indefinite retention. Policy can reclaim machine exhaust or sensitive records; non-sensitive tombstones can preserve references where appropriate. They can’t preserve the missing explanation. The history must make that loss crystal clear.

Recovery and plumbing should be ordinary

The staging area should disappear from the user model entirely. We’ve been complaining about it for years. The next implementation will very likely still need an index, cache, or materialization database, but we as developers shouldn’t have to manage it as a second place where changes temporarily live.

Jujutsu proves the DX win already. It snapshots working-copy changes when relevant commands run, records repo mutations in the operation log, and lets jj undo undo the last operation without asking the user to reconstruct ref movement. JJ does this very well. This protects recorded repo state, but it can’t recover uncaptured edits or undo external effects. jj split and jj squash move work between changes without turning the index into a user-facing state.

When repository mutation is recoverable, users can rewrite, split, reorder, and combine work without treating every history edit as dangerous. We will all have less anxiety around collaborative editing. Automatic snapshots, explicit change selection, change-editing operations, and ordinary undo are better primitives. Safe actions should be obvious; recovery shouldn’t require a reflog excavation.

IDEs, agents, CI planners, and forges need a stable library or protocol for querying identities, deps, content, conflicts, operations, and capabilities. They should definitely not have to scrape CLI output or reverse-engineer a mutable working directory. Exposing that model through a library or protocol also lets the CLI stay small (another gripe of mine with respect to Git is the massive command surface).

Migrate from Git; don’t design around it

I don’t want Git compatibility dictating a successor’s design. This is a mistake. This is dragging legacy concepts around in a world where they’re not actually needed anymore. A successor that requires every critical repo and integration to abandon Git on day one is still dead on arrival… but we don’t have to choose compatibility. We can choose migration and we can treat it with the respect it deserves.

Migration likely needs to import Git objects, refs, authorship, and history, and provide Git-compatible endpoints or adapters so teams can migrate workflow by workflow. Existing Git signatures require retaining the original signed payloads for verification; copying a signature onto a translated revision doesn’t authenticate that revision. A migration engine can move history. It can’t, by conversion alone, keep Git-only integrations working.

Stable change identity, committed conflicts, project memory, asset policies, and richer dependencies may not round-trip through Git. An export must preserve that state separately, reject the operation, or require an acknowledged loss. Silent flattening isn’t compatibility and will not cut it.

Git is today’s standard, so the migration path will decide whether teams can adopt a successor. I’d likely build the migration engine alongside the VCS, or even before, and make it excellent: teams will need help bringing their history and workflows across; it will need adapters only where and while they still need them. I don’t want the native design carrying Git’s constraints after those constraints have stopped being useful. This is why we’re having this conversation to begin with.

Where the VCS stops and the forge begins

All four requirements depend on what survives the loss of a forge. The VCS should own that project history: material, changes, revisions, dependencies, conflicts, and enough provenance to explain why a state was accepted or rejected. This is non-negotiable from my perspective.

The forge owns live production coordination: review policy, execution, distribution, deployment, organization identity and capabilities, and external integrations. Shared identities connect them. Recording that an engineer or agent (belonging to the engineer) produced a revision belongs in durable history; deciding what that engineer, or agent, may deploy today doesn’t.

This “barrier” is also why I think a VCS controlled by one company has to be open source end to end, server included. The financial incentive to build a great product is real - Paul Graham recently outlined this really well in a video on X/Twitter - and a company can charge for hosting, scale, and operations. But if the history stays fully interpretable only while that company’s server exists, the VCS has quietly become part of the forge. This isn’t the direction the community should go in. It will not improve VCS through competition; it will create confusion and more fragmentation. The VCS is the one piece of software where more competition is a bad thing, in my opinion, in as much as they’re closed source in any way at all.

Unfortunately, Git is good enough - until a successor is much better

“Good enough” explains both Git’s dominance and the lack of serious appetite for replacing it. To say the engineering lift is enormous is a gross understatement. The ecosystem cost is likely even worse. Every incremental improvement makes the immediate pain easier to tolerate and gives us another reason to put off replacing it. We’ve watched Cursor and Cloudflare deploy real engineering hours and capital to maintain the Git protocol rather than fix it. This is a warning to all of us. The road is treacherous.

Putting off innovation here has worked while the missing state lived in our heads. In fact, if we weren’t all using agentic workflows today - this wouldn’t even be talked about outside of a few nerds in separate corners of the world. Agents are doing more of the work now, and they start every session without it. This isn’t the future. Git will be another victim of accelerated progress, and that is a good thing.

Of course, any successor has to justify that work while leaving room for at least the four requirements above… and likely many more I’m not thinking of right now.

Git won. It deserved to win. It still needs to be replaced, and whatever replaces it will have to make all of us willing to accept the migration. The silver lining is the migration will be much easier for the same reason it’s necessary - the tools we have today (LLMs, harnesses, etc.) will make it much more manageable.

Coming next: Part Two asks why replacing GitHub is harder still. Its systems have to move together.

loadingalias

engineering diary

About

I’m loadingalias — a Rust engineer and startup founder. I build infra that makes complex software smaller, faster, and easier to use.

The thread through all of my work is collapsing unnecessary complexity and legacy fragmentation. I care about systems with fewer moving parts, fewer competing sources of truth, explicit proof boundaries, portable implementations, and world-class performance. I target modern hardware in everything I do. I aim to keep things bound by only physics and my vision.

This is where I write - incredibly opinionated - articles on Rust, databases, distributed systems, dev tooling, cryptography, version control systems, storage engines, and the engineering decisions behind my work.

Connect