Formatting a 25M-line codebase overnight
143 points
by r00k
8 hours ago
| 15 comments
| stripe.dev
| HN
CrzyLngPwd
7 hours ago
[-]
One of my first jobs was a small software company writing software for a small number of clients, in MS basic PDS.

The lead developer didn't like to bother with formatting code, so I wrote a tool called makenice to format his nasty spaghetti gibberish into something with good indents and layout to make it easier for us normal people to parse.

He was furious, literally spun in circles about it right in the office in front of everyone, so I wrote makenasty to format code into the way he appeared to like.

I only shared makenasty/nice with a couple of the team, who loved it, as it allowed easy conversion between something readable and something the team lead like.

He never knew about makenasty.

reply
munk-a
6 hours ago
[-]
Outside of the naming - this is a perfectly sane thing to do for developer comfort and can usually be accomplished with simple transformations.

There are often limitations (like manually added indentation/spacing for alignment) but as long as you're very intentional about what changes you'll allow and have a good understanding of the language it can be an extremely safe operation.

reply
nitwit005
6 hours ago
[-]
If he didn't bother formatting code, it would seem impossible to create a tool that formatted code the way he preferred.
reply
singpolyma3
6 hours ago
[-]
Sounds like he did format code, and even had opinions on how it should be formatted, but OP disagreed.
reply
Terr_
6 hours ago
[-]
I find a lot of these conflicts I can't resolve when everybody agrees that the pain of ugly/unnecessary diffs is greater than the pain of minor formatting disagreements.
reply
jameson
5 hours ago
[-]
reminds me of rob pike mentioning gofmt's style is "no one's favorite"
reply
ethical_source
5 hours ago
[-]
This kind of passive-aggressive bullshit is exactly what's wrong with tech. People don't decide things: they just passively resist, and authority ends up being a muddle of truncated information flows.
reply
munificent
7 hours ago
[-]
> We chose a Saturday to format the entire codebase to avoid merge conflicts. And while our test suite gave us high confidence we'd gotten everything right, it's always a bit daunting to have a diff so large that GitHub can't render it.

The dart formatter has an internal sanity check. It walks through the unformatted and formatted strings in parallel skipping any whitespace. If any non-whitespace characters don't match, it immediately aborts. This ensures that the only thing the formatter changes is whitespace, and makes it much less spooky to run it blind on a huge codebase.

That sanity check has saved my ass a couple of times when weird bugs crept in, usually around unusual combinations of language features around new syntax.

(Unfortunately, the formatter in the past year has gotten a little more flexible about the kinds of changes it makes, including sometimes moving comments relatively to commas and brackets, so this sanity check skips some punctuation characters too, making it a little less reliable.)

reply
Terr_
6 hours ago
[-]
I imagine a fancier version would be to compare the Abstract Syntax Trees.
reply
caminanteblanco
2 hours ago
[-]
The only issue is then you're at the mercy of whatever parser your formatter uses to construct the AST
reply
hyperhello
50 minutes ago
[-]
Strictly speaking that wouldn’t work, since a1 is different from a 1, for example.
reply
hobofan
7 hours ago
[-]
I'm surprised they went with a all-at-once reformat. Even when doing it over a weekend this is bound to mess with a lot of open PRs at their scale.

I had to introduce a formatter in a few sizeable codebases in the past (few 100k to few million LOC), and I always did it incrementally via a script that reformatted all files that are not touched in any open PR. The initial run reformatted 95% of all files. Then I ran the script every day for ~two weeks and got up to 99.5% of all files and then manually each time one of the remaining ~dozen PRs that were WIP for longer were merged.

reply
rileymichael
6 hours ago
[-]
both options have their pros and cons. if you utilize some form of ratcheting[1], you can sneak it in without your team knowing.. but all of your PRs for the foreseeable future will have a ton of reformatting screwing with your git blame. if you do it all at once, someone will have to sort out conflicts, but you can utilize `blame.ignoreRevsFile`[2] so that your history remains useful

[1] https://github.com/diffplug/spotless/tree/main/plugin-gradle...

[2] https://git-scm.com/docs/git-blame#Documentation/git-blame.t...

reply
hobofan
6 hours ago
[-]
Yes, that is a good point. This is also why I personally would recommend to let a central person/team handle the reformatting rather than sneaking it into every PR (- see my sibling comment). That way you can be in charge of having a uniform style of commit messages to make the reformat commits easy to identify and create a well kept ignoreRevsFile. I think that provides the best of both worlds.
reply
BobbyTables2
6 hours ago
[-]
That’s a neat feature, thanks for sharing.

Unfortunately I find that code bases lacking auto formatting are often littered with non functional changes as developers temporarily instrument code, remove it, but leave whitespace changes behind.

In terms of tracking code changes, one really would have to rewrite the entire history with each commit reformatted.

reply
WhyNotHugo
3 hours ago
[-]
> I'm surprised they went with a all-at-once reformat. Even when doing it over a weekend this is bound to mess with a lot of open PRs at their scale.

Rebasing PRs should be trivial, just rewrite all commits reformatting the files it touches, then rebase, `git checkout --theirs`, run formatter again, and `git rebase --continue`. It's methodical and scriptable, you don't need to manually resolve any conflicts.

reply
skydhash
7 hours ago
[-]
You can always let the team know so that they can apply the formatter on their PR branch.
reply
hobofan
6 hours ago
[-]
In the smaller migrations I did I tried that, but some way or another a decent chunk of the people still managed to get stuck in merge/rebase conflicts. I would almost explicitly not recommend giving that advise to the teams.

My rough blueprint for introducing formatter or linter nowadys would be:

- Recorded knowledge share session around how to set up the tools for local use 1-2 weeks before the initial rollout, and outline how the process will take place

- On the day of the initial rollout send out a reminder + the recording again

- Do the initial PR

- Incrementally do the rest of the migration, and subscribe to the PRs that drag out the process

reply
jrajav
6 hours ago
[-]
This is exactly the remedy to the PR issue. I've "lucked" into owning a Prettier formatting pass at two different places now, and did the same process at each - full pass on master, simple step-by-step process to follow to update any PR by running the format script.
reply
zx8080
4 hours ago
[-]
It's probably because the author can shit on others (let me guess, a senior principal something engineer).
reply
varun_ch
8 hours ago
[-]
I’m shocked at the 25M line part! That is a completely unfathomable amount of code for one codebase. I really want to know more about that.
reply
phoyd
5 hours ago
[-]
I am more shocked by the "overnight" aspect. I tried running clang-format on the Chromium source (68,281 .cc files, 21 million lines according to wc):

$ find chromium-149.0.7826.1/ -name ".cc" -exec cat {} + | wc 21640925 55715244 833460441

And that took less than 6 minutes on a single E5-2696 v3 from 2014:

$ time find chromium-149.0.7826.1/ -name *.cc | parallel -j 16 clang-format $x>/dev/null

real 0m5.666s user 1m13.964s sys 0m13.373s

That’s orders of magnitude faster, especially if we assume they’re not running their workloads on potatoes like mine. Is Ruby’s syntax really that much more complicated than C++, or is this a tooling problem?

reply
deets87
5 hours ago
[-]
I don't think the post necessarily means it took multiple hours to format the codebase, I think they're probably just saying they worked on it off-hours and landed it while no one was working so that it didn't run into merge conflicts.
reply
christophilus
5 hours ago
[-]
My guess would be tooling. I think the Ruby formatters are written in Ruby. I’d guess the clang one is written in C.
reply
bruckie
7 hours ago
[-]
Only 25 million? :) Google had billions a decade ago...

https://research.google/pubs/why-google-stores-billions-of-l...

reply
Groxx
5 hours ago
[-]
iirc they also vendor(ed) many of their dependencies, several layers deep, which still counts for "stores" though it's rather different than "wrote" / "maintains".
reply
bruckie
1 hour ago
[-]
Very true. It was still hundreds of millions of lines of first party code a decade ago, and could easily be over a billion at this point.
reply
Groxx
13 minutes ago
[-]
Yeah, I can definitely believe that Google would break over a billion handwritten. It's a big company that has been around for a long time.

It's still absurd. But believable.

reply
jsnell
7 hours ago
[-]
Right, where is the rest of the code?
reply
mr_mitm
7 hours ago
[-]
They're up to 42 million now, as per the article
reply
lukan
7 hours ago
[-]
That sounds even more insane to me, but I guess most of that code does not really touch financial transactions, otherwise it would be a nightmare being responsible to verify that.
reply
clintonb
6 hours ago
[-]
Ruby code touches financial transactions. Card payments were migrated to Java when I left in 2022. Non-card payments (e.g., ACH, checks, various wallets) were still processed by Ruby.

PCI-related/vaulting code lived in its own locked-down repo. I think that was a mix of Go and Ruby.

Once you have the foundations in place for account balances and the ledger, processing a payment isn’t that daunting. Those foundations, however, took a lot to build and evolve.

reply
jamesfinlayson
3 hours ago
[-]
> Once you have the foundations in place for account balances and the ledger, processing a payment isn’t that daunting. Those foundations, however, took a lot to build and evolve.

Pretty much. I've worked at places with PHP payment processing that worked just fine, and at a place with C++ payment processing (and no testers) and it worked just fine. I wasn't around when the systems were first built though so not sure if there were tears along the way.

reply
varun_ch
4 hours ago
[-]
> migrated to Java

I want to know more about this

reply
deathanatos
5 hours ago
[-]
My (much smaller than Stripe) company is well over 4.5M at this point, and the graph is very much exponential.

AI has been a huge problem here: the amount of code is just exploding. Quality of the produced code is another matter.

reply
Neywiny
5 hours ago
[-]
^^^^^^^^^^^^^^^^^^^

I recently wrote a very esoteric Python script. 100 lines of code. No classes, no functions, but yes argparse.

I've tried out the latest open source models on the task. They go bananas. It's like Enterprise fizzbuzz (https://github.com/enterprisequalitycoding/fizzbuzzenterpris...). They love classes and imports and reinventing the wheel. A great way for me to tell trash AI slop code is it'll define a useful constant then 15 lines later do it again with a different name.

They love making code that looks impressive. "Wow look at all the classes and functions. It's so scalable. It's so dynamic. It validates every minutae against multiple schema and solves a problem I never thought about." But it was trash code. One really was 400 lines and it didn't even look like it would work. Can't even imagine what it means for 4.5M moderately good human lines to become what? 27M fluffy filler repeat lines that don't even make sense?

reply
manoDev
5 hours ago
[-]
The bad part of LLM is it got trained on bad examples because us humans also don't know WTF we're doing.
reply
Neywiny
3 hours ago
[-]
Yeah maybe I need to do the old "you are a veteran engineer" nonsense. I've had some success telling it to implement everything it suggests and be production ready. I hate when it takes a shortcut and says I'll have to change it. That's kinda the whole point of me not writing the code...
reply
dgrin91
3 hours ago
[-]
I don't understand why the felt the need to do a big-bang merge like this. Its a formatter, so the files should be functionally equivalent before and after. Why not just enable it for new files/edit files for a while, then once comfortable apply it to old files in batches? What advantage does the big bang merge give? Seems higher risk for the same reward
reply
fsckboy
2 hours ago
[-]
could introduce subtle bugs, so doing it all at once while it's on the front of everybody's mind with as much comprehensive review and testing of parts or the whole to everybody's satisfaction. if you don't do it all at once, you'd need to repeat the same amount of testing multiple times.

>files should be functionally equivalent before and after

when you say something like this, the road you are on is paved with good intentions.

reply
eigenblake
2 hours ago
[-]
Really reminds me that there's nothing in principle stopping us from storing parse trees and exposing them via something git like so we can avoid even needing to format, let alone also needing to resolve a whole category of merge conflicts based on that formatting. I mean a format is just a theme over your data -- I mean code.
reply
nitwit005
6 hours ago
[-]
> Given that complexity, the hypothesis was simple: tackle the hardest syntax first and the rest will follow.

Always nice to see. I've seen people fall into the trap of designing for the common case, not realizing most of the code will be to deal with the less common cases.

reply
sgc
1 hour ago
[-]
In another field I have heard it called going for the jugular; the vivid description helps get the point across nicely.

If you want to master something, you will have to know the hardest part. So just deal with that first and then everything else is easy, because you are dealing with it as somebody who has already mastered the domain.

reply
burnte
7 hours ago
[-]
The floating spiral thing is so distracting I spent more time deleting it in Inspector than reading the article. I feel like they hate their readers. Awful.
reply
annaspies
6 hours ago
[-]
If you set `prefers-reduced-motion: reduce`, it goes away
reply
tmaly
3 hours ago
[-]
How did I know this was going to be a rewrite in Rust?
reply
comrade1234
6 hours ago
[-]
Man must me nice to have the time to put so much work into tabs.
reply
Pxtl
2 hours ago
[-]
Clean indenting is about saving time so you don't spend way too long getting lost trying to understand what seems like an insane piece of code until you realize it was a mundane bug hidden by incoherent indentation.
reply
ryanisnan
1 hour ago
[-]
Cool story. The treat at the end was fun as well, thank you!
reply
hokkos
7 hours ago
[-]
Now it makes me wonder, are those 45M LoC are untyped ?
reply
c3ab8ff137
7 hours ago
[-]
No, Stripe has its own Ruby typechecker - https://sorbet.org/
reply
m12k
7 hours ago
[-]
reply
CrzyLngPwd
7 hours ago
[-]
Surely, it no longer needs to be human-readable, and the era of write-only code is finally upon us with the dawn of AI writing our mealtickets.

Why bother formatting 25m lines of slop, and why is AI wasting tokens on making code look human-readable anyway?

reply
sgc
1 hour ago
[-]
Every LLM I have ever asked about this says they perform better when they receive pretty-printed code because it is easier to see structure and priorities. It has been an almost universal recommendation for me, and it makes sense since LLMs are just mimicking human expression.
reply
throwatdem12311
3 hours ago
[-]
What is even the point of formatting code anymore.
reply
cadamsdotcom
6 hours ago
[-]
An insight about code is that compared to the scale we operate on data, code as text is tiny. Instantaneous git operations and “run this tool over all the code” are the norm even while we wait for LLMs to stream their tokens to stream back so tool calls can operate on it.

That insight might seem obvious - but if you stay cognizant of it as you work, you can invent some pretty amazing tooling for yourself & your team.

reply