Top
Best
New

Posted by ibobev 1 day ago

The GitHub wiki is an anti-pattern (2022)(michaelheap.com)
172 points | 109 comments
gwking 1 day ago|
The last paragraph says: > At some point your docs will outgrow a single folder, and then all bets are off. You’ll want a separate repo with its own build process...

My question is, why is this taken as a given? Is it so hard to have docs and code live together in version control after a certain scale? If so, what is the specific problem and what is the cause?

I ask because I've never been that satisfied with the various ways I've tried to organize projects in git. Recently I've been trying to keep the source, tests and docs together in the same tree so that changes are more localized. It seems to be helping me keep track of things, especially with coding agents so eager to make changes all over the place. I find their proclivity to repeat the same idea in multiple locations (agent instructions, docs, docstrings, help strings, comments) especially problematic.

bluGill 1 day ago||
On a large project you will have problems. You can maintain a monorepo anyway as many people do, and deal with the problems of a large monorepo. Or you can go to multirepo and deal with the issues of multirepo. Both have been done successfully, and both have significant problems that you need to work with.

Most people advocating a monorepo have never worked on a project large enough to see the issues with a monorepo and so are arguing for a monorepo without understanding the problems with them. For most people a monorepo is the correct answer because their project is small.

dualvariable 1 day ago|||
Seems like if you're small enough, a monorepo is the right way to go because it doesn't matter at that scale, and if you're big enough, you'll have the resources to throw at making monorepos scale.
bluGill 1 day ago||
Mono vs poly at scale needs resources. You have different compromises with each and so the resources go to different places. However there is no clear cut winner despite a few mono repo at scale advocates trying to claim otherwise - they are always completely ignoring the issues with a monorepo setup.
dualvariable 1 day ago|||
Multirepo at scale also needs resources, there's an enormous amount of work required for version bumping and synchronizing everything. People always completely ignore all those chores. Generally, you have either a massive amount of tech debt or you have one person doing nothing but running around doing all that work for everyone else (works great if all that work seems to magically appear for you). You can also be furiously working at automating all those chores, but that's the same level of effort you'd have to throw at scaling out a monorepo--just different.
bluGill 1 day ago||
That was my point - there are issues with both at scale. Trying to pretend that one is better for everyone is wrong. You just choose your tradeoffs.
sshine 23 hours ago|||
Exactly because monorepos have least overhead when they’re small, monorepos generally win because you need to be small for a long while until you get big.

By the time you’re “at scale” (who knows), and all these monorepo at scale problems start to overwhelm, you can switch strategy, because the economy of polyrepos is so obvious by then.

So far, I’ve started a new job a handful of times by collapsing a premature polyrepo strategy: people were not experienced enough to merge two git repos without a common root.

I’ve only once went the other way, and it incurred so much overhead, it decreased developer productivity by some small but not insignificant percentage.

To be clear: I’m not a maximalist. All of my open-source work is exceedingly compartmentalised. My DNS library is separate from my external-dns webhook is separate from my fork of external-dns. They could all live in one repo. But FOSS encourages reusability, commercial software encourages clumping and vendoring.

gilfaethwy 21 hours ago||
> By the time you’re “at scale” (who knows), and all these monorepo at scale problems start to overwhelm, you can switch strategy, because the economy of polyrepos is so obvious by then.

Conversely to your experience, I have worked at a handful of places who have a monorepo that has been creaking under its own weight for years, but its structure as a monorepo now underpins the business, and so migration to a polyrepo simply never happens, and developers are now checking out a 50GB repo in its entirety periodically.

horsawlarway 19 hours ago||
I'm on the other side of this problem, with a company that went multi repo for bad reasons (political, not technical) and I would give you serious money if you could solve my problems by just forcing me to check out 50gb every now and then...

Instead I deal with a 30+ repo clusterfuck (technically we have 60+ services, but I only have to run half...) that is held together by hopes and prayers, takes literal hours of actual effort to bring everything up to date on master, and has become a fractured hellscape where people are afraid to leave their tightly constrained silos of service combinations.

Long story short... I will take a bad monorepo over bad multirepo any day of the week.

bluGill 18 hours ago|||
I suspect that your hours of effort to get everything up to date would exist in a monorevel, too. It would exist in a different form, and so it would be harder to measure, but a large part of the work has to be done either way. It's just that certain parts of the work become very visible.
dualvariable 14 hours ago|||
Yeah, compared with the maintenance effort of running around to 100 different repos, keeping everything in sync and the deps all updated, I'd always take the pain of 50GB checkouts.
dualvariable 58 minutes ago||
It looks like these days a combination of git-scalar, git-lfs, and bazel will work for a small 50GB monorepo like that. The largest cost would be porting over the existing build system to bazel and then integrating it with CI. Once that is done, though, the monorepo will just scale, and all the constant ongoing cost of multirepo sync will be gone.
cortesoft 1 day ago|||
> Most people advocating a monorepo have never worked on a project large enough to see the issues with a monorepo

Really? I feel like most of the stuff I have read advocating monorepos are from people at Google, which is a HUGE monorepo.

ibobev 20 hours ago|||
Isn't Microsoft also using a monorepo?
bluGill 1 day ago|||
Google is an advocate of monorepo. However a lot of people are seem to be regurgitating what google wrote about them, but they are not Google scale and have no idea what the problems Google faces are.

Google also is very much in the yell loudly and ignore anyone who points out the problems of a monorepo.

cortesoft 21 hours ago||
Maybe it is just because I know a lot of Googlers/ex-Googlers, but almost all of the advocates for monorepos that I have spoken to are basing it from their experience at Google. I won't disagree that they tend to gloss over the extensive tooling they have to make it work, though...
johannes1234321 1 day ago|||
> Is it so hard to have docs and code live together in version control

Managers and others won't touch the repo. (Sometimes it's better the don't...)

ragall 22 hours ago|||
We had managers and even non-engineers check-in documentation at Google: each page had an "Edit me" button that spawned an editor in a new tab with a CL(PR) ready. It worked very well.
giancarlostoro 1 day ago|||
> I find their proclivity to repeat the same idea in multiple locations (agent instructions, docs, docstrings, help strings, comments) especially problematic.

I have not taken full advantage yet, but every source sub-directory can have its own CLAUDE.md or AGENTS.md file, with instructions for a given directory, whenever something like Claude opens a file in a folder, if there's an appropriate agent doc in the housing folder, it should read / apply its contents when working. Not sure how this happens with Sub-Agents on that note.

I think if you want both, you might as well add a ./docs/ directory, and put your md documents in there however you want, this has the upside of letting you have docs with your code always, as well as letting you link to direct source code files.

darau1 1 day ago|||
Weirdo checking in

I always initialize my projects with a src, and docs, directory, for exactly this reason.

My reasoning is that I shouldn't have to go hunting for the docs for the code, or vice versa.

guipsp 1 day ago||
Do you have only those two dirs at the top level? If so how are you finding it? I tend to have a docs dir at the top level, along with other build stuff
darau1 23 hours ago||
I've had no trouble so far. I'd have no problem with ephemeral build artifacts going either at the root, or in src.

edit: I have had one pain point: devs/managers ask me why I do that, and that I stop. I refuse.

TheRoque 1 day ago||
I also put everything in one repo nowadays, no matter the usage, the language etc. All is synced, and all is accessible by my LLM. There are a lot of tools to manage monorepos, and frankly most of the time you don't even need them.
ffsm8 1 day ago|||
It makes releasing software a lot more complicated.

Not a big deal if you're basically the only developer and handroll the process - but poly repositories make release and dependency management a lot more straightforward to wrangle

bluGill 1 day ago|||
Most of the time your repo is so small that you won't run into the problems of a large repo and so you don't need those tools. Don't confuse that for monorepos have no problems when they get large.
genxy 1 day ago||
How large is large? What does large mean? More than 100GB? 2TB?
bluGill 1 day ago||
You have the wrong measure. Large is number of parts. That is partially source files, partially projects/teams, and partially things that are conceptually not related. Likely other things as well.

It is unlikely GB/TB is ever a measure, though if your repo is that big and some people only need a subset of the repo it would be.

jonjon10002 23 hours ago||
Worked at a place (as a tech writing manager) where the monorepo was several TB and company-issued laptops all had 500 GB hard drives. It was like a rite of passage that new writers or people with new laptops would inevitably not RTFM or learn about sparse checkout and try to clone the entire repo, which was not good.
genxy 22 hours ago||
Nice hazing ritual, I hope that behavior extended to other parts of the organization.
chungy 1 day ago||
Fossil (https://fossil-scm.org/home/doc/trunk/www/index.wiki) solves this pretty nicely. You can have documentation as files or in a special wiki namespace and it's versioned both ways, and every repository clone gets everything. Even better than that, your in-tree documentation files are rendered and browseable in exactly the same way as the dedicated wiki namespace.

The linked URL to the home page there can even serve as an example: the "trunk" is a check-in name (https://fossil-scm.org/home/doc/trunk/www/checkin_names.wiki) that points to the newest check-in on the "trunk" branch. You can replace it with any other reference to get the old version; eg, version-2.20 would work to get the version 2.20 of the docs, 2015-03-14 would work to get the version from 14 March 2015, etc.

YPCrumble 1 day ago||
Why is this easier or more effective than just a /docs directory?
chungy 1 day ago|||
You absolutely can use "just a /docs" directory in Fossil. You can even point the web server to /doc/trunk/docs/index.md or whatever other file names you want. :-)
gatlin 1 day ago|||
Parent comment linked to that answer.
mghackerlady 1 day ago||
Fossil is the best. Sqlite uses it
bigfishrunning 1 day ago|||
Fossil was written for Sqlite in the same way that git was written for Linux. It's really a shame that more projects don't use it. I think that a github competitor (with social features, PRs, CI, etc) with a fossil backend would be very popular.
somat 21 hours ago||
In fossil's case, every instance includes those features, this makes the social platform "the web itself"

Having said that, there are central hosting projects, but where that has some value with git, git proper is distributed and does not need it but all the auxiliary stuff does. with fossil it provides almost no value. All the auxiliary stuff is built in so all it provides is a place to host. Which is fair but hosting is not hard with fossil.

https://chiselapp.com/

With everything included in fossil I abuse it as a personal social platform(think discord) super easy to host and it gets me chat, forums, wiki, and file storage. None of them great, but it is so easy I don't really care. voice and video do require an additional service, so there is that. Now I just need to find some friends...

rpdillon 1 day ago|||
Yep, I use Fossil for all my side projects. Super easy to host, tiny, includes everything I need for a project, all in one file. Great piece of software.
yellowapple 21 hours ago||
I've been gradually migrating my Git repos to Fossil as I've been touching them. Been quite happy with it.

There are only a couple things that I miss:

- Grouping repos together into a combined project. I'd love to be able to have a single Fossil server/instance with a single set of users, wiki pages, tickets, etc. but multiple independent codebases. Closest I've gotten to that is to simply have multiple independent branches instead of a single trunk (example: https://fsl.yellowapple.us/avorion/home), and it's worked surprisingly well, but it's clear Fossil wasn't designed with this workflow in mind, so there are some rough edges with it (albeit minor and easy to work around).

- Compatibility with things that assume you're using Git. Being able to export to Git helps a lot here, but it's still extra steps. Ideal solution here would be for the Fossil server to be able to double as a Git forge and present repos accordingly, such that I can point things like CI/CD pipelines or Terragrunt module calls or whatever directly to the Fossil repo itself over the Git interface those things expect instead of having to setup a Git forge manually and somehow synchronize everything w.r.t. access controls.

- An equivalent to Git's submodules. This would IMO help address the “grouping repos together into a combined project” case as well.

somat 20 hours ago||
There are login-group options which lets you have a single sign on with a group of repos, To unify the other bits it looks like it would be possible to change the main menu links to the central repo(in the web settings).

Never wanted to do what you want but it sounds possible

set up the fossil server on a directory of repos probably with --repolist to get a list

    fossil server --repolist /var/fossil/
get all repos in the same login-group

change the menus in each repo to point to the core repo. Based on some quick tests I suspect a full http:// url is required.

    - Forum     /forum       {@2 3 4 5 6}
    + Forum     https://mydomain.org/core/forum       {@2 3 4 5 6}
mikeocool 1 day ago||
In my experience, the docs for something like setting up a dev env are typically greatly improved by the second person who sets up the dev env, not the personal who originally wrote the docs.

In that case, when the docs are not associated with a code change, you want to make getting those improvements into the docs as frictionless as possible, otherwise the changes aren't going to get made.

Personally, I've found that making docs updates incredibly fast + easy to be far more valuable than anything you get from forcing doc changes through the full SDLC process. If someone has feedback on your docs changes they would have shared in a review, they can just update the docs instead.

yunwal 1 day ago||
> In my experience, the docs for something like setting up a dev env are typically greatly improved by the second person who sets up the dev env, not the personal who originally wrote the docs.

In my experience, this is also true of a lot of code as well. Your dev scripts should probably have much more relaxed standards than your service source or CI/CD. Ideally I could define merge requirements by directory without doing some weird shenanigans with the CODEOWNERS file and a bot.

juancn 1 day ago|||
That could be easily be corrected by relaxing merge gates for changes only to the `docs` folder (or some suitable naming pattern).

You can even do live edits on the web if you don't want to use a command line.

wavemode 1 day ago|||
You can set up automation and/or configuration such that changes to the docs folder don't require code review.
jameshart 1 day ago||
Corrections and improvements to docs are just a bugfix though?
codazoda 1 day ago||
I'm no fan of GitHub add-ons and I agree with the premise here but...

I can think of one other possibility. It's easier to write in a wiki via the browser. I can open that on my phone and edit docs. I can open it in my browser and edit docs.

On desktop it's a tiny bit more to pull the repo and open it in your editor (and you might already be there) but that tiny bit can be enough to stop you from writing documentation. For me, writing documentation must be totally painless so that I'll actually do it.

Why am I not a fan of the add-ons like PR's, wiki's, discussions, projects, and issues? Because they each introduce vendor lock-in to varying degrees.

sheept 1 day ago||
Wikis don’t have as much vendor lock-in as other Github features since they’re just git repos,[0] so you can clone and push the wiki elsewhere.

[0]: https://docs.github.com/en/communities/documenting-your-proj...

wky 1 day ago|||
GitHub is pretty good about editing on the web. Markdown files can be edited straight from github.com. On desktop you can hit the period key to directly open the repo in vscode.dev. Technically on mobile you can change github.com to github.dev to do the same, though the editing experience is worse than directly editing on GitHub.
cocoto 1 day ago||
You can edit single files in most git forges and it will open a pull request for you, the workflow is not that bad.
ericyd 1 day ago||
I disagree, requiring code review for docs changes sounds great but in my experience it's extremely hard to get a human to review docs changes. Either you get a rubber stamp with no real review (zero added value, adds useless friction) or you spend days bugging people to actually review your changes. All for docs!

The counter-argument i envision is: "update your docs and code at the same time in the same PR!" That works great, until you want to document something that isn't precisely tied to a single piece of code. In fact I think the most useful docs describe high level systems rather than being associated with specific pieces of code. Use comments for that; in contrast, docs should be easily editable by anyone at all times, otherwise they never get updated (an evergreen problem in any scenario).

solatic 10 hours ago||
This is one of the arguments for polyrepo: different sources have different sensitivities. Code ending up in production needs to be reviewed, so enforce reviews.

Docs do not.

So what results is either allowing developers to push (but not force push) directly to main (you can always push revert commits if needed), or a PR process that exists to enforce linters and build-ability, but if those pass, allow the developer to merge independently, without human review.

anon7000 17 hours ago|||
Agreed, but I think the review & CI system should ideally be able to ignore markdown changes. Easier said than done.

But I think the biggest benefit, which we shouldn’t overstate, is that agents will just update docs as they find them. Including for big picture systems. Keeping docs updated is a PITA.

Writing style of AI often sucks. But I’ve found it pretty easy to rectify. And having some correct context is better than nothing or outdated docs in a lot of cases.

yellowapple 21 hours ago||
This'll be a controversial answer, but this sounds like exactly the sort of thing an LLM should be able to do reasonably well, whether by the human writing the docs and the LLM updating the code accordingly or by the human writing the code and the LLM updating the docs accordingly.
stephenlf 1 day ago||
I agree with this post. I’ve never found the GitHub wiki experience to be particularly ergonomic. I don’t have any issues with it, but it’s no more convenient than a simple /docs folder. And from there, it’s almost trivial to turn /docs into GitHub pages. Similar effort for a much better end product.

Wikis typically connote distributed, anonymous edits. This feature is partially covered by git already.

cxr 1 day ago|
> I’ve never found the GitHub wiki experience to be particularly ergonomic.

That's because the original sin of GitHub "wikis" is that they weren't (and most of them still aren't) even wikis. There's this perverse thing that happened during the wiki age, where people unable or unwilling to get on board decided to just start calling things "wikis" even though they exemplify the very thing that the wiki was invented as a response to. The reckless debasing of the word then infected adjacent spaces. Sourcehut's "read-only wikis" (wat) aren't even designed to be edited in the browser; on Sourcehut, "Publishing your changes is as easy as committing them and pushing them upstream." Newsflash: That's not a wiki.

masklinn 1 day ago||
Yeah you can configure gh “wikis” to be freely editable but that’s not the default and most of them are not,
joknoll 1 hour ago||
Those GitHub Wikis are also just repos. Instead of cloning the repo.git you can clone repo.wiki.git. And then combined with git subtrees you have the best of both worlds. A local docs folder in your main repo and a browser github wiki.
sigvef 1 day ago||
Not many people know that the github wiki is actually backed by a separate "hidden" repo, and can be accessed by adding .wiki to the repo url, e.g. https://github.com/lionleaf/dwitter.wiki.git

Apparently it can even do CI stuff. Still wouldn't recommend it though, for the other reasons outlined in TFA and this thread.

isityettime 23 hours ago||
It's in the docs, so it's not exactly hidden:

> You can edit wikis directly on GitHub, or you can edit wiki files locally.

https://docs.github.com/en/communities/documenting-your-proj...

> Every wiki provides an easy way to clone its contents down to your computer. Once you've created an initial page on GitHub, you can clone the repository to your computer with the provided URL.

Every GitHub wiki you visit also has a section at the bottom that points to that URL with the description

> Clone this wiki locally

as well.

WorldMaker 1 day ago|||
I seem to recall that GHE allowed you at one point to point a Wiki to a /docs or /docs/wiki subdirectory. I still don't know why that never became a public GitHub.com feature. Sure it is technically redundant with the file browser if you know how to use the view controls and don't mind the UI sprawl of the file browser, but I still think it would be a good feature.
annex-winged-cr 1 day ago||
It leads to a 404 page on my end.
sheept 1 day ago||
It’s not a web page; it’s a git remote url
a4isms 1 day ago||
The second paragraph neatly triggered my confirmation bias:

The initial version of this post opened with “You can use the wiki or a docs folder for your GitHub project, both are valid choices” but as I wrote more, I realised that there is a single reason to use a wiki, and many more reasons not to use the wiki. So many in fact, that I consider using the wiki on GitHub is an anti-pattern.

A very straightforward example of McCulloch's quote that "Writing is thinking."

swiftcoder 1 day ago|
I think there is a missing pro here on the wiki side: trivial edits are trivial. Even fixing a typo in the docs directory requires PR + approvals + CI. Effectively limiting your docs contributors to folks who are comfortable with a code editor, and git, is a decision
Kinrany 1 day ago|
Forcing CI and approvals on changes to docs/ is a decision
swiftcoder 1 day ago||
Entertainingly, excluding a directory from PR approvals is not an option GitHub provides out of the box
pocksuppet 1 day ago||
Then don't do approvals. Comment "approved" and if someone merges without someone else commenting "approved" someone gets mad at them.
swiftcoder 1 day ago||
One can always work around tools limitations. The limitations tell us something about how the creators intended it to be held, however
pocksuppet 31 minutes ago||
You don't have to hold a tool the way its creator intended. Do you call yourself a hacker?
More comments...