Secrets Management for DevOps Teams
.env file, somebody else uses the same key in another application, and a year later nobody knows where the key is used or who owns it. Good secrets management is not only about storing passwords in a vault. It is about managing the full lifecycle of a credential: how it is created, where it is stored, who can access it, how long it lives, how it is rotated, and what happens when it is compromised.
For DevOps teams, I would start with one principle:
If you can avoid having a secret at all, do it.
Don’t put secrets into code
This sounds obvious, but secrets still regularly end up in Git repositories.
For example:
DB_PASSWORD = "ProductionPassword123!"
API_KEY = "abc123-production-api-key"
The same applies to configuration files:
database:
username: app
password: ProductionPassword123!
Moving a password from source code into a YAML file does not make it secure if the YAML file is committed to Git.
The same is true for Base64. Base64 is encoding, not encryption.
And a private repository should not be considered a secret store either.
Repositories get cloned to developer computers, copied into CI/CD systems, backed up and integrated with third-party tools. Even after a secret is removed from the current version of a file, it may still exist in Git history.
The better approach is to keep only a reference in the application and retrieve the actual secret when it is needed.
db_password = get_secret("production/database/password")
The application knows which secret it needs. It does not contain the secret itself.
Before storing a secret, check if you actually need one
This is probably the most important part.
Instead of asking:
Where should we securely store this access key?
ask:
Why do we need a permanent access key in the first place?
A lot of cloud and CI/CD authentication can now work without long-lived credentials.
For example, a pipeline may traditionally have something like:
AWS_ACCESS_KEY_ID
AWS_SECRET_ACCESS_KEY
AZURE_CLIENT_SECRET
Those credentials may stay valid for months or even years.
A much better model is workload identity or federation.
CI/CD pipeline
↓
Workload identity
↓
OIDC / federation
↓
Cloud IAM
↓
Temporary credentials
↓
Deployment
GitHub Actions, GitLab, Azure DevOps and modern cloud platforms support variations of this model.
The pipeline proves who it is, and the cloud platform gives it temporary access.
There is no permanent cloud password sitting inside the CI/CD system.
The same idea should be used for applications where possible. Prefer managed identities, IAM roles, workload identities and short-lived tokens instead of permanent API keys.
My preference is roughly:
Workload identity → temporary credential → automatically rotated secret → static long-lived secret
The further right you go, the more risk and operational work you create.
Use a proper secrets manager
Some secrets cannot be removed. In that case, store them centrally.
Depending on the environment, this could be Azure Key Vault, AWS Secrets Manager, Google Secret Manager, HashiCorp Vault or another enterprise secrets-management solution.
The exact product is less important than the operating model.
You should be able to answer:
- Where is this secret stored?
- Which application uses it?
- Who can read or modify it?
- When was it created?
- When was it last rotated?
- Can I revoke it quickly?
- Who owns it?
If nobody can answer these questions, then the secret is not really under control.
Avoid spreading credentials across Git repositories, CI/CD variables, Kubernetes YAML files, spreadsheets, documentation pages and developer laptops.
The more copies you have, the harder rotation and incident response become.
One credential should not be used everywhere
A very common shortcut is to create one powerful service account or API key and use it across several applications.
It is convenient until it leaks.
If ten applications use the same credential, compromise of one application potentially compromises all ten. It also becomes difficult to understand where the leak came from.
Use separate identities for separate systems, environments and workloads.
Development and production should not share credentials.
Application A and Application B should not normally use the same service account.
A read-only process should not receive write or administrative permissions.
The principle is simple:
Give the identity only the access it needs, to the resources it needs, for as long as it needs it.
CI/CD needs special attention
CI/CD systems are especially sensitive because they often have access to several important things at once:
- source code
- container registries
- cloud environments
- Kubernetes clusters
- infrastructure-as-code
- production deployments
- application secrets
A compromised pipeline can therefore become a direct path into production.
Production credentials should only be available to jobs that actually need them.
Pull requests, forks, test pipelines and untrusted branches should not automatically receive production secrets.
Where possible, use OIDC federation instead of putting static cloud credentials into the CI/CD platform.
If a static secret is unavoidable, use the platform’s protected secret storage and restrict who can create, update and use it.
Also remember that hiding a value behind ****** in the pipeline log does not make the credential secure.
If someone can modify the pipeline, they may still be able to send the secret somewhere else.
This is why access to modify CI/CD configuration should itself be treated as privileged access.
Be very careful with logs
A secret can be stored perfectly and still leak through logs.
For example:
Connecting to database
server=db01
username=production-app
password=SuperSecretPassword
Now the secret is potentially sitting in your SIEM, logging platform, troubleshooting exports, support tickets and backups.
The same problem happens with:
- environment-variable dumps
- authorization headers
- connection strings
- debug output
- command-line arguments
- Terraform output
- exception messages
- CI/CD debugging logs
Sensitive values should be masked or excluded before they are logged.
Logging systems often have much broader access than production secret stores, so this should not be underestimated.
Environment variables are useful, but they are not a vault
Using an environment variable is better than hardcoding:
DATABASE_PASSWORD=${DB_PASSWORD}
But the important question is still:
Where did DB_PASSWORD come from?
If somebody copied it manually from a spreadsheet into a pipeline three years ago, you still have a static unmanaged credential.
Environment variables are a way of delivering configuration to an application. They are not, by themselves, a secrets-management solution.
Ideally, the secret comes from a vault or the application authenticates using its workload identity.
Also be careful with debugging and support tools because environment variables can sometimes end up in diagnostic output.
Kubernetes Secrets are not automatically secure
Kubernetes has a resource called Secret, but the name can sometimes create a false sense of security.
A Base64-encoded Kubernetes Secret is not encrypted simply because the value looks unreadable.
If Kubernetes Secrets are used, make sure you have proper RBAC, encryption at rest, restricted cluster administration and auditing.
For more mature environments, integrate Kubernetes with an external secrets manager or use workload identity.
If a Pod can authenticate directly using its runtime identity, that is generally better than injecting permanent cloud credentials into it.
Infrastructure-as-Code can leak secrets too
Terraform, Helm, Ansible, Bicep, CloudFormation and similar tools should reference secrets instead of containing them.
Terraform deserves particular attention because sensitive information can appear in the state file.
Even if the Terraform code contains:
sensitive = true
this mainly controls how values are displayed. It does not necessarily mean that the sensitive value does not exist in the state.
Terraform state should therefore be protected like any other sensitive asset:
- use a protected remote backend
- encrypt it
- restrict access
- audit access
- do not put state files into Git
The same principle applies to generated deployment files and build artifacts.
Rotation should not break production
Security teams often say that credentials should be rotated regularly.
That is correct, but rotation needs to be designed properly.
If changing a database password causes an outage because the application still has the old password, then the rotation process itself becomes a production risk.
Where possible, support overlapping credentials:
Credential V1 → currently used
Credential V2 → created
Credential V2 → distributed
Credential V2 → tested
Credential V1 → revoked
This allows credentials to be replaced without downtime.
For important machine credentials, rotation should ideally be automated.
More importantly, make sure you can perform emergency rotation.
If an API key leaks today and replacing it requires three days, five approvals and a maintenance window, you have an incident-response problem.
Short-lived credentials are better than rotation calendars
A credential that automatically expires after an hour is usually much safer than one that stays valid until somebody remembers to rotate it.
Expiration reduces the value of a stolen credential.
This does not mean every secret needs a five-minute lifetime. The right lifetime depends on the system.
But permanent credentials should become the exception.
This is also where machine credentials and human passwords should not be confused.
For service accounts and application credentials, automatic rotation and short lifetimes make sense.
For human accounts, forcing people to change passwords every 30 or 60 days is a different issue and often creates predictable passwords instead.
For users, focus more on unique passwords, password managers, MFA or passkeys, compromised-password detection and immediate password changes when compromise is suspected.
Stop secrets before they reach Git
People make mistakes.
Someone will eventually paste a key into a configuration file and try to commit it.
The solution should not depend only on training.
Use technical controls.
Secret scanning can be implemented at several levels:
Developer workstation
↓
Pre-commit scanning
↓
Git push
↓
Push protection
↓
Pull request / CI scanning
↓
Continuous repository scanning
The earlier you detect the secret, the easier the problem is to fix.
Pre-commit tools can catch credentials locally.
CI pipelines can scan commits and pull requests.
GitHub and similar platforms can detect known secret formats and, in some cases, block the push completely.
Run secret scanning against private repositories too. Private does not mean safe.
If your organization uses proprietary tokens or credential formats, add custom detection patterns where possible.
Deleting a leaked key from Git is not enough
One of the most important operational rules is:
If a real secret has been committed to Git, consider it compromised.
Deleting it in the next commit does not solve the problem.
The secret may still exist in Git history, developer clones, build systems, backups or third-party integrations.
If the repository was publicly accessible, assume automated systems may already have found it.
The first action should be to revoke or rotate the credential.
Cleaning Git history can still be useful, but it comes after revocation.
Do not waste valuable time trying to prove that nobody saw the key before rotating it.
Monitor how secrets are used
Secrets management should not stop at storage.
You should also monitor how sensitive identities and credentials are being used.
Useful detections include:
- a service account authenticating from an unusual location
- production credentials being used from a developer workstation
- unusual numbers of secret retrievals
- a workload accessing resources it has never accessed before
- repeated failures using an expired credential
- human users reading production secrets unexpectedly
- credentials being used after an application has been decommissioned
Secret-store logs, cloud IAM logs, CI/CD audit logs and application logs should be available for investigation and, where appropriate, connected to the SIEM.
During an incident, you should be able to answer:
Who accessed the credential, when did they access it, and what did they do with it?
Have an owner
A secret without an owner will eventually become an old secret that nobody wants to delete.
For important credentials, keep at least basic ownership information.
| Field | Example |
|---|---|
| Secret | Payment API credential |
| Owner | Payments Team |
| Environment | Production |
| Used by | payment-service |
| Permissions | Submit and query payments |
| Storage | Secrets manager |
| Rotation | Automatic, every 30 days |
| Monitoring | Enabled |
| Long-term plan | Replace with workload identity |
This information can live inside the secrets-management platform, CMDB or another inventory system. It does not need to be a spreadsheet.
The important thing is that somebody owns the credential and is responsible for its lifecycle.
Have a clear process for leaked credentials
When a secret leaks, do not start by discussing whether the attacker probably saw it.
Start with containment.
- Identify the credential.
- Understand what access it provides.
- Revoke or rotate it.
- Check where and how it was used.
- Look for suspicious activity.
- Find every place where the credential was stored.
- Remove unnecessary copies.
- Fix the process that allowed the leak.
The last point is important.
If a developer commits an API key and the only response is “please be more careful next time,” the same thing will happen again.
If the organization enables secret scanning and push protection afterwards, the incident has resulted in an actual security improvement.
A practical baseline for DevOps teams
If I had to reduce everything above to a short organizational standard, I would use these rules:
- Never store secrets in source code.
- Never treat a private Git repository as a secret store.
- Prefer workload identities and temporary credentials over static keys.
- Put unavoidable secrets into an approved secrets-management platform.
- Use separate credentials for separate workloads and environments.
- Apply least privilege.
- Do not print secrets into logs.
- Protect Terraform state and other deployment artifacts.
- Enable secret scanning and push protection.
- Automate rotation where possible.
- Make sure important credentials can be rotated quickly during an incident.
- Treat leaked credentials as compromised.
- Monitor sensitive credential usage.
- Assign an owner to important secrets.
- Keep reducing the number of permanent credentials in the environment.
Final thought
Secrets management is sometimes treated as a storage problem:
Where should we keep our passwords and API keys?
I think the better question is:
Why do we have this secret, and can we remove it?
Vaults are important. Rotation is important. Secret scanning is important.
But the best outcome is often to replace the secret completely with an identity-based authentication mechanism.
Instead of this:
Application
↓
Password stored somewhere
↓
Permanent access
↓
Resource
try to move toward this:
Application
↓
Verified workload identity
↓
Authorization policy
↓
Short-lived access
↓
Resource
That removes a whole category of problems around storage, copying, rotation and accidental leakage.
For DevOps teams, I would summarize secure secrets management in five points:
Don’t hardcode secrets. Reduce the number of secrets. Prefer identities over passwords. Give minimum access. Automate the lifecycle.