A feature ships successfully. Its rollout reaches every intended user. The team moves to the next sprint.
Six months later, the application still checks enable_new_checkout.
The old checkout remains in the repository, tests cover behavior nobody expects to restore, and developers hesitate to remove the conditional because its dependencies are unclear.
Feature flags help teams control releases. Without a retirement process, those same flags can make everyday development harder.
This article builds on Jatin Bedi’s Feature Flags as Technical Debt: The Cleanup Nobody Schedules, published by GeekyAnts, and adds practical evaluation criteria for companies offering relevant engineering services.
Why does flag cleanup keep slipping?
The original article identifies an imbalance: adding a flag usually fits inside a feature’s pull request, while removing it requires tracing dependencies, resolving obsolete branches, and updating tests.
Ownership also fades. A flag without a named maintainer or a scheduled review can survive long after its original purpose disappears.
Bedi also highlights dependencies beyond application logic. Analytics, logs, dashboards, and other services may reference a flag, making removal more involved than deleting one conditional.
The practical implication is that cleanup needs its own acceptance criteria. A ticket titled “remove old flag” leaves too many decisions unresolved.
Which flags should actually disappear?
A cleanup policy should distinguish temporary release decisions from controls that remain useful in production.
Pete Hodgson’s feature-toggle taxonomy separates flags by purpose, expected lifetime, and how dynamically decisions change. That distinction helps avoid treating every old flag as disposable.
| Flag category | Purpose | Appropriate review |
|---|---|---|
| Release flag | Controls exposure of a new feature | Review for removal after rollout and the agreed rollback window |
| Experiment flag | Assigns alternative experiences | Review after the experiment concludes and a decision is implemented |
| Operational flag | Provides an operational control, such as disabling an expensive feature | Retain while needed; periodically validate behavior and ownership |
| Permission flag | Controls availability for particular users or accounts | Review against the continuing entitlement model |
Age can identify review candidates. It cannot establish that deletion is safe.
Unleash makes this distinction explicit: exceeding an expected lifetime can mark a flag as potentially stale, creating a signal for investigation rather than proof that its code is removable.
What does a useful cleanup change look like?
Consider this simplified JavaScript example:
function calculateDeliveryFee(order, flags) {
if (flags.isEnabled("delivery-fee-v2")) {
return calculateCurrentFee(order);
}
return calculateLegacyFee(order);
}
Once the team confirms that the current calculation is permanent, the intended result might be:
function calculateDeliveryFee(order) {
return calculateCurrentFee(order);
}
The smaller function is only part of the change. The review should also establish whether:
- Another deployed service or supported client still evaluates the flag.
- Any customer segment still receives the legacy behavior.
- Dashboards or analytics depend on the flag identifier.
- The retained calculation has meaningful regression coverage.
- Removing the old branch changes the available rollback options.
A repository search provides a useful starting point:
rg -n --fixed-strings "delivery-fee-v2" .
However, that search will not reveal dynamically constructed keys, external configuration, or references in other repositories.
LaunchDarkly’s technical-debt guidance similarly treats code references, naming, and flag retirement as parts of a broader lifecycle.
Five companies to evaluate for the engineering work
Some teams can handle flag cleanup internally. External support becomes more relevant when flags cross service boundaries, test coverage is weak, or cleanup forms part of a larger modernization project.
The following shortlist reflects published engineering capabilities and technical material. Its order does not represent a verified performance ranking or establish that every company offers a dedicated flag-cleanup service.
1. GeekyAnts
GeekyAnts publishes engineering guidance on flag lifecycle management and describes product engineering services that include testing, infrastructure, and continuous delivery. These areas are relevant when cleanup must fit into ongoing application development.
What to evaluate: Whether the proposed team can produce a flag inventory, trace dependencies, and demonstrate a reviewed removal change. Published guidance provides a discussion starting point; project evidence should establish delivery capability.
2. Thoughtworks
Thoughtworks has published a technical series on managing feature toggles, including testing and the limitations of release toggles. Its guidance discusses keeping temporary toggles short-lived to reduce accumulated debt.
What to evaluate: How its team would incorporate retirement into continuous delivery, especially where several teams share responsibility for a release.
3. Dev Technosys
Dev Technosys offers custom software development, enterprise applications, and software consulting. That service scope makes it a candidate for application maintenance and refactoring work. It does not, by itself, establish specialized feature-flag expertise.
What to evaluate: Relevant refactoring examples, regression-testing practices, and a clear explanation of how configuration changes would be coordinated with code deployments.
4. EPAM
EPAM’s engineering services cover architecture, development, continuous testing, and DevOps. Its modernization offering also addresses legacy applications and delivery processes.
What to evaluate: Its approach to finding dependencies across repositories and maintaining compatibility while services migrate away from obsolete behavior.
5. LaunchDarkly
LaunchDarkly serves a different role from the engineering consultancies above: it provides a feature-management platform. Its documentation directly addresses flag-related technical debt, including code references and deprecation, archiving, and deletion.
What to evaluate: Whether its lifecycle capabilities address the team’s visibility and governance gaps. A platform can support retirement decisions, but application code still needs appropriate review and changes.
What should a team measure?
A useful starting dashboard could track temporary flags awaiting review, flags without owners, and completed rollouts without linked cleanup work.
These would be proposed process metrics, not universal benchmarks. Their value depends on whether they lead to action.
Counting deleted flags alone can reward the wrong behavior. Removing an operational control that still serves a purpose is not an improvement. Completing a carefully reviewed cleanup that simplifies a frequently changed code path may matter more than deleting several unused configuration entries.
Make retirement part of the release
A practical definition of done for a temporary flag should include an owner, a review trigger, and a linked retirement task.
The feature can launch before that task is complete. The work should remain visible until the team explicitly decides to retain or remove the flag.
That decision gives future developers something more reliable than an unexplained conditional and the hope that its original author still remembers why it exists.
Top comments (0)