How to Start Working on an Unfamiliar Codebase
You have access to the repository.
Now comes the less obvious part: working out how to actually work on it.
On an established team, the code is only one piece of the development environment. There are runtime versions, build tools, supporting services, secrets, test data, IDE settings and conventions that everybody else may have stopped thinking about years ago.
Your first goal is not to understand the entire codebase.
It is to get from:
I have cloned the repository
to:
I can build it, run it, test it, debug it and make a small change.
Once that loop works, learning the code becomes much easier.
Look for the bootstrap before you build anything
Before installing a single thing, find out whether somebody has already automated the setup.
Look for:
- A
READMEat the root of the repository, especially a setup or development section - A
CONTRIBUTINGfile, which often contains details the README skips - A
scripts/ortools/directory containing setup or bootstrap scripts - A
.devcontainerdirectory - A
docker-compose.ymlorcompose.yml - A
Makefile,Taskfileor similar - An onboarding page on the team's wiki
Do not just read the first few paragraphs of the README and start improvising.
I once spent three days assembling a development environment by hand from a wiki page and a lot of guessing. On day four I found a bootstrap script in the repository.
It had been there for two years.
It was linked from a README I had only read the top half of.
If a bootstrap process exists, start there even if you would have designed it differently.
If it fails, that is useful information. You have found a problem in the onboarding path that the next person will probably hit too.
Find the versions the project expects
A surprisingly common cause of "it works for everyone but me" is a version difference.
Not necessarily a dramatic one.
You might have Java 21 while the project expects Java 17, a different Node version, or a package manager release that behaves slightly differently.
Most ecosystems have some way of declaring the expected version:
- Node:
.nvmrc, anenginesfield inpackage.json, or avoltasection - Python:
.python-version, or a version constraint inpyproject.toml - Java: a Gradle toolchain or settings in
pom.xml - .NET:
global.json - Ruby:
.ruby-version - Multiple runtimes:
.tool-versionsif the team uses asdf or mise
Use those before installing whatever version happens to be newest.
A version manager is useful here. Tools such as nvm, fnm, pyenv, SDKMAN, rbenv, asdf and mise let different projects use different runtime versions without constantly reinstalling them globally.
If the documentation and the project configuration disagree, look at what is actually building the software today.
The CI pipeline is often very useful for this.
Check .github/workflows, azure-pipelines.yml, .gitlab-ci.yml or whatever your team uses.
Look at:
- which runtime versions it installs
- which package manager it uses
- which commands it runs
- which order those commands run in
CI is not necessarily the complete description of the local development environment, but it is a useful example of a configuration that is currently expected to build and test the project successfully.
Find out what the application needs around it
Your application probably does not run on its own.
It may depend on:
- a database
- a cache
- a message queue
- another internal service
- a local emulator
- a shared development environment
- test data
- API credentials
- a private package or artifact repository
Find out how the team runs those dependencies.
Sometimes everything comes up through Docker Compose. Sometimes developers run a database locally but connect to shared versions of other services. Sometimes most of the environment already exists elsewhere and the application just connects to it.
Two questions are particularly useful early on.
Where does the data come from?
An empty database may let the application start while still making it impossible to do useful development.
Look for a seed command, migrations, fixtures, an anonymised database dump, or another documented way of creating development data.
If you cannot find one, ask.
Where do the dependencies come from?
Many companies host their own packages, libraries and build artefacts in a private repository such as JFrog Artifactory, Sonatype Nexus, Azure Artifacts, or GitHub Packages.
Your project may therefore be downloading dependencies from an internal repository rather than directly from npm, Maven Central, PyPI, NuGet, or another public registry.
Companies use internal registries so they can control which packages and build artefacts their teams depend on. An internal registry can cache public packages, host private company libraries, restrict unapproved dependencies, scan packages for security issues, and keep builds working even if an external registry is unavailable. It also gives the organisation a single place to manage access and see what software is being used across its projects.
This is worth checking early. If dependency installation fails, the problem may not be your package manager at all. You may simply be missing access to the company's artifact repository, or the project may need additional repository configuration or credentials.
Look at the project's package manager and build configuration to see which repositories it expects to use.
Where do the secrets come from?
Your local environment may need connection strings, API keys or credentials for development services.
Those should normally come from a documented process, secrets manager, vault or managed configuration.
They should not come from a colleague pasting production-looking credentials into a chat message.
You may also find a local environment file such as .env.
Check whether there is a .env.example or similar file showing which variables are required without containing the real values.
If there is not one, that may be worth fixing later.
Do not personalise the project yet
There is a temptation on a new machine to immediately recreate the setup you had at your last job.
Your shell, editor, aliases, formatter, linting rules and automation.
Some of that is harmless.
Bring the things that only affect you:
- Your shell and prompt
- Your aliases
- Your editor keybindings and colour scheme
- Your terminal and window management
- Personal productivity tools that do not modify the repository
Be more cautious with anything that changes project files.
My first pull request at one job contained about twenty lines of real work inside a diff of several hundred lines.
My editor had reformatted two files automatically using settings the project did not use.
The reviewer was polite about it.
He also, reasonably, asked me to come back when he could see what I had actually changed.
If the project has a formatter configuration, use it.
If it has an .editorconfig, respect it.
Use the project's linting rules rather than your global ones.
Be careful with import organisers, automatic refactoring tools and anything else that silently rewrites files.
A useful rule is:
Anything that changes the repository is a team concern.
Anything that changes only your screen is yours.
Once you understand the project conventions, you can decide which of your usual tools fit safely.
Prove the whole development loop works
The application starting is not the end of setup.
You are ready to work when you have completed the whole development loop at least once:
- Clone the repository
- Install the dependencies
- Build the project
- Run the tests
- Run the application
- Make a trivial change and see it take effect
- Attach a debugger and stop on a breakpoint
- Commit the change to a branch and push it
Do this before you are under pressure to deliver something real.
Debugger setup in particular is easy to postpone because the application appears to work without it.
That becomes painful the first time you actually need to investigate a problem.
Get it working while the stakes are zero.
Pay attention to the test suite as well.
If tests fail before you have changed anything, write down which ones.
Otherwise, a week later you will make a change, see a red test and spend an hour trying to work out whether you broke it.
Follow one real piece of behaviour through the code
Once the project runs, resist another temptation: trying to understand the entire repository.
You probably cannot, and you do not need to.
Instead, choose one small piece of behaviour and trace it through the system.
For a web application, you might start with one endpoint.
Find:
request
↓
controller or handler
↓
business logic
↓
database or external service
↓
response
For a frontend application, choose one screen or interaction and follow where its data comes from and what happens when the user does something.
The goal is not to memorise every layer.
You are learning how this particular codebase is put together.
As you follow one path, you will naturally discover where the application starts, how modules are organised, where configuration lives, how dependencies are wired together, how data is accessed and where tests tend to sit.
That gives you a much better mental model than reading directories from top to bottom.
Make one harmless change
Once you understand one path through the application, change something small.
It does not need to be useful.
Change a piece of text. Add a log statement. Adjust something in a local-only path. Make a tiny modification that lets you prove that you understand how a source-code change becomes running software.
Then:
edit
↓
build
↓
test
↓
run
↓
debug
If that works, you now have a functioning development loop.
That is a much more useful milestone than simply getting the application to start.
Keep a setup log from the first command
Open a file called SETUP.md before you start and write things down as you go.
Not afterwards.
As you go.
Record:
- commands you needed
- settings you changed
- errors you hit
- what fixed them
- documentation that was wrong
- steps that were missing
- things somebody had to tell you
You will almost certainly need some of this again.
You may get a new machine, work on another project, rebuild a corrupted environment or help the next developer through the same setup.
More importantly, you are in a useful position right now.
You can see every gap in the onboarding process because none of it is obvious to you yet.
In a few weeks, much of it will feel obvious and you will forget which parts were confusing.
That makes your setup notes useful raw material for improving the team's actual documentation.
How much this varies
Some teams give you a dev container or a bootstrap command and ten minutes later everything works.
If that is your situation, appreciate it. Somebody put real effort into making that possible.
Other teams have a setup process that mostly lives in one person's head. The documentation is a two-year-old Slack thread and the only reliable instructions begin with "ask Sam."
In that environment, the job is partly archaeology.
Read the repository.
Read the pipeline.
Read the Compose files.
Ask specific questions.
Write down the answers.
The most common situation is somewhere in between.
There is documentation, but part of it is stale. A setup script exists, but one step has rotted. A service has been renamed. A version has moved on. Something that used to be manual is now automated.
Those problems can make excellent early contributions.
You have just experienced the setup process as a new developer, which means you are unusually well placed to see where it fails.
Then there is the fully remote development environment, where most of the work happens in a cloud workspace or virtual desktop.
The details are different, but the goal is the same.
Get to the point where you can confidently build, run, test, debug and change the software.
Once you can do that, you are no longer just looking at somebody else's codebase.
You can start working in it.