Private Package Registries: What They Are and Why Companies Run Them
Most package managers default to fetching dependencies from a public registry. npm talks to registry.npmjs.org, pip talks to PyPI, Maven resolves from Maven Central, and so on. In many companies, that is not what happens. An internal private registry sits in the way, and your tools need to be pointed at it.
If you have ever seen a 401, a 403, or a "package not found" error for something that clearly exists on the public registry, this is often why.
The common products are JFrog Artifactory, Sonatype Nexus, Azure Artifacts, GitHub Packages, and AWS CodeArtifact. They all solve the same problem with the same underlying concepts. The examples below use Artifactory for specifics, but the reasoning applies to whichever product is in use.
Why companies do this
It looks like bureaucracy from the outside. It is mostly not.
- Public registries have outages and packages get removed. Caching dependencies internally means a public outage does not stop a build.
- An internal registry gives you one place to track what is being pulled in, block vulnerable packages, and answer "are we using the thing that was just on the news?" in minutes rather than days.
- Some open source licences carry obligations a company cannot accept commercially. Catching that at the point a dependency is added is far cheaper than discovering it during an acquisition.
- Reconstructing exactly which dependency versions shipped with any release is a genuine requirement in regulated industries, and useful everywhere.
- Internal libraries, container images, and build outputs need a home that is not a public registry. The artifact repository provides that too.
The three concepts that make it make sense
A company needs to do two different things with packages: store the ones it builds itself, and safely access the ones the outside world builds. These are kept separate, then merged at a single URL for convenience. Every artifact repository product has this same three-tier model, though the names vary. In Artifactory they are local, remote and virtual. In Nexus they are hosted, proxy and group.
The internal store (local or hosted) is where your company publishes its own packages: shared libraries, build artifacts, container images. Think of it as a private file store. Nothing from outside the company writes here.
The caching proxy (remote or proxy) is where most of the value is. This is a server the company runs on its internal network - behind the corporate firewall, not reachable from the public internet. When you install a package, your machine talks to that internal server rather than going directly to npm or PyPI. The first time anyone requests a package, the server fetches it from the public registry and stores a copy. Everyone afterwards gets the cached copy from inside the network. This is what keeps builds running when a public registry has an outage, and what gives the company a single chokepoint to inspect or block incoming packages.
The aggregated view (virtual or group) solves the practical problem: you do not want to configure two separate URLs everywhere, one for internal packages and one for the public cache. This tier merges them behind a single URL. Your package manager points at one place and never needs to know whether a package came from internal storage or the cached public registry.
The practical consequence: you should almost always be pointing at the aggregated endpoint, and you usually need exactly one URL per ecosystem. If somebody has given you three URLs to configure for one package manager, ask whether a virtual or group repository exists, because it probably does.
Pointing each ecosystem at it
Each package manager has its own configuration file and its own way of carrying credentials. The URL shape varies by product, so take the exact URL from your internal documentation rather than from the examples below.
npm, yarn and pnpm
In your user level .npmrc:
registry=https://artifactory.example.internal/artifactory/api/npm/npm-virtual/
//artifactory.example.internal/artifactory/api/npm/npm-virtual/:_authToken=${NPM_TOKEN}
Note the second line's odd syntax: the registry URL without the scheme, then a colon, then the credential field. This is npm's format for per registry credentials and it is easy to get subtly wrong. A missing or extra trailing slash will produce a 401 that looks exactly like a bad password.
I lost an hour of my life to a trailing slash in that second line. The registry URL in my credential line did not exactly match the URL npm was requesting, so the token was never applied, and every response came back as a 401. I spent that hour regenerating tokens and becoming increasingly convinced my account was broken, when the credential had been correct all along and was just being attached to the wrong address.
Scoped registries are worth knowing about. You can send only your company's own scope internally and let everything else go to the public registry:
@example:registry=https://artifactory.example.internal/artifactory/api/npm/npm-internal/
Whether your organisation wants that or wants everything through the virtual repository is a policy decision. Follow the team's convention.
NuGet and .NET
A nuget.config, either in the repository or at user level:
<configuration>
<packageSources>
<clear />
<add key="internal" value="https://artifactory.example.internal/artifactory/api/nuget/nuget-virtual/index.json" />
</packageSources>
</configuration>
The <clear /> matters. Without it, your source is added alongside the public one, and restore may resolve from either. That is the mechanism behind dependency confusion attacks, where a package published publicly under your internal name gets picked up instead of yours. Clearing the defaults and listing exactly one source is the safe pattern.
Maven and Gradle
Maven uses settings.xml in your .m2 directory, with a mirrors section pointing at the virtual repository and a servers section carrying the credentials. Gradle takes repository definitions in the build script, and credentials from gradle.properties at user level rather than in the repository.
The rule for both: the URL can be committed, the credentials cannot. Repository definitions belong in the project, credentials belong in your user level configuration where version control cannot reach them.
Python
[global]
index-url = https://artifactory.example.internal/artifactory/api/pypi/pypi-virtual/simple
Credentials can be embedded in that URL, which is convenient and means your password ends up in a config file, so prefer a token and a restrictive file permission, or a keyring if your environment supports one.
Be careful with extra-index-url. Unlike index-url, it adds a source rather than replacing, and pip may resolve from either, which is the same dependency confusion risk as the NuGet case.
Also resist --trusted-host as a fix for certificate errors. That is the certificate problem from the previous article and it deserves the certificate solution, not a flag that disables verification.
Docker and container images
docker login artifactory.example.internal
Images are then referenced with the internal registry as a prefix. Note that docker login writes credentials to ~/.docker/config.json, base64 encoded and not encrypted, unless a credential helper is configured. Base64 is encoding, not protection. On a shared or backed up machine, that file matters.
CLI tools
Most artifact repository products ship a CLI: jf for Artifactory, nx for Nexus, and platform CLIs for cloud-hosted options. These can configure credentials for you, upload and download artifacts, and publish build information that links a release to exactly the dependencies it used.
For day to day development you often do not need them. They become worth learning when you are working on the build itself, publishing artifacts, or when your organisation requires build traceability.
Credentials, and where they belong
Do not use your password. Most artifact repositories issue identity tokens, API keys or reference tokens, which are revocable, scoped and safe to rotate. Some integrate with single sign on and can generate a token for you from the web interface.
Where credentials go:
- User level configuration files, outside any repository.
~/.npmrc,~/.m2/settings.xml,~/.gradle/gradle.properties - A credential manager or keyring, where the tool supports it
- An environment variable populated from one of the above
Where they must never go:
- A configuration file inside the repository, even for a moment. Secret scanners will find it, and a token in Git history stays there
- A Dockerfile, or a container image layer
- A screenshot, a chat message or a wiki page
- The build pipeline as a literal value rather than a stored secret
Two more habits worth forming. Tokens expire, so note the date, because a build that has worked for months failing on a Tuesday morning is often an expired token. And CI should use a service account rather than your personal credentials, since a pipeline authenticating as you will break the day you change role or leave, and it makes the audit trail meaningless.
Reading the errors properly
The failures in this area are unhelpfully similar. Learning to tell them apart is most of the skill.
404or "package not found" for something that clearly exists: you are pointed at a repository that does not include the public source, or at a local repository rather than a virtual one.401 Unauthorised: credentials are missing, malformed, or expired. Check the exact registry URL in your credential line, including the trailing slash, because a mismatch means the credential is never applied.403 Forbidden: you are authenticated but not permitted. This is a permissions issue, not a configuration problem.403with a policy message about a licence or vulnerability: a security policy block from the repository's scanning layer (JFrog Xray, Nexus Lifecycle, or similar). The package exists, but this one is refused. Find out which policy triggered rather than looking for a workaround. The answer is usually a newer version of the package.A hang followed by a timeout: almost always the proxy. The internal registry should not be going through it.
A TLS error: the registry is presenting a corporate certificate that the tool does not yet trust.
The first time I hit one of these I assumed the registry was misconfigured and went looking for somebody to fix it. It was not misconfigured. A package two levels down my dependency tree had a known vulnerability, the policy had done exactly what it was written to do, and the answer was a newer version that had already been released a fortnight earlier.
Those last two are worth spelling out, because this is where the proxy and certificate articles collide with this one. Your artifact repository is an internal host. It belongs in NO_PROXY, and it will present your corporate root certificate. Getting either wrong produces symptoms that look like an authentication problem and are not.
The lockfile trap
One thing that catches teams out rather than individuals.
Lockfiles in several ecosystems record the URL each package was resolved from. If you install using the internal registry, your lockfile may now contain internal URLs. That is fine internally, and it is a problem if the repository is ever open sourced or shared with a partner, and it is confusing for anyone resolving from a different registry.
Check what the team's convention is before your first dependency change. Some organisations configure this away, some strip it in CI, and some accept it. Either way, it is worth knowing before you change a lockfile.
How much this varies
- Azure Artifacts and GitHub Packages integrate with their own platforms, so authentication looks different, but the same three-tier model still applies. AWS CodeArtifact issues short-lived tokens you fetch with a CLI command, which surprises people the first time a build fails twelve hours in.
- Many companies use public registries directly with no internal repository. Faster to start, less to configure; the trade-offs are the ones listed above.
- In air-gapped environments the internal registry is the only source of packages, populated deliberately, and adding a new dependency is a formal request with a review. That is slow by design, a response to a real threat model rather than an obstacle for its own sake.
If you are stuck right now
- Confirm you have the right URL, and that it is a virtual repository rather than a local or remote one.
- Confirm the registry host is in your proxy bypass list.
- Confirm the corporate certificate is trusted by the tool you are using, not just by your browser.
- Confirm your credential is a token rather than a password, is not expired, and is attached to the exact URL your tool is requesting.
- Read the status code rather than the message. 401 is you, 403 is permissions or policy, 404 is the wrong repository, a timeout is the proxy.
Then write down what fixed it, because the documentation for this is nearly always out of date.