Skip to content

Practical software architecture

Design Principles for Software That Survives Change

Software architecture should let a business evolve without repeatedly rebuilding its platform. These eight principles come from years of building enterprise commerce systems with small teams, changing vendors, and production constraints. They guide decisions about boundaries, reliability, and the cost of future change. Each carries a tradeoff: the useful test is whether a decision makes the next necessary change safer, less expensive, and easier for another engineer to understand.

Principle 1

Build Around Business Capabilities

Back to principles

Business capabilities should outlive the vendors and technologies that implement them. Keep core workflows expressed in business terms so replacing a provider does not require rewriting the rules around it.

In practice: Checkout asks a payment capability to authorize an amount. An adapter translates that request into the provider's API and returns an outcome checkout understands, including an indeterminate result.

Tradeoff: Adapters add mapping and contract tests. They must preserve meaningful provider differences, including unsupported operations. Use them where replacement is plausible; direct SDK usage can be reasonable for a short-lived experiment.

Read the MongoDB-to-PostgreSQL migration case study for the boundaries that contained change and the operational work they could not remove.

Explore the reasoning: Build Around Business Capabilities

Keep vendor details at the boundary

An authorizePayment contract can accept an amount, currency, payment method reference, order reference, and idempotency key. Its outcomes should distinguish approved, declined, indeterminate, and failed requests. Vendor status codes, request objects, and retry rules belong in the adapter.

A generic execute(data): any interface hides a vendor without explaining the operation. Copying every SDK method into a new interface preserves the coupling under another name. Model the capability callers need, including its failure behavior.

The same approach applies to other integrations:

  • A Shopify adapter can expose catalog publication, order retrieval, fulfillment updates, and refunds without spreading webhook or GraphQL response shapes through orchestration code.

  • A loyalty interface can describe balance lookup, earn, redeem, and reversal while retaining provider identifiers and error details at the boundary.

  • A repository can expose findOrderByExternalId or savePromotionRedemption while keeping storage representation and query details local.

Preserve differences that affect the business

Providers are not interchangeable in every respect. If a payment provider cannot separate authorization from capture, represent that limitation or reject the unsupported workflow. Keep provider transaction identifiers available for reconciliation and support.

Repository boundaries should reflect business aggregates and transaction behavior. A universal storage interface that treats SQL, document storage, caches, and search engines as identical usually removes useful capabilities. Abstract where replacement is valuable, without building a private database framework.

Principle 2

Optimize for Change

Back to principles

Design for changes the business has evidence it will make. Put likely variation behind clear boundaries while keeping the stable workflow straightforward. Flexibility is useful when it lowers the cost of a real next step.

In practice: Keep checkout's sequence stable while tax calculation, promotion eligibility, and fulfillment selection vary through explicit policies. Start with one payment adapter; add routing when a second provider creates a concrete need.

Tradeoff: Every extension point adds another concept to maintain. When a rule is still changing shape, ordinary code and a later refactor can cost less than preserving an early guess.

Explore the reasoning: Optimize for Change

Choose boundaries from recurring changes

Pricing rules, fulfillment sources, marketing campaigns, and provider APIs change repeatedly in commerce. Identify those recurring decisions before introducing extension mechanisms. A likely change deserves a clear boundary; a hypothetical future does not automatically deserve a framework.

Choose the mechanism to fit the variation:

  • Use configuration for understood, bounded changes that are safe without deployment.

  • Use composition when small, testable operations can assemble the required behavior.

  • Use adapters at external boundaries.

  • Keep ordinary code when the shape of the rule is still uncertain.

A promotion pipeline can compose a few explicit rule types without becoming a general-purpose expression language. Introduce more flexibility when the second rule or implementation provides evidence of what must vary.

Include the operational migration

Replacing a vendor involves more than its API client. Historical identifiers, reporting, reconciliation, webhooks, retries, and support tools must continue to work. Plan data backfills, parallel operation, cutover, and rollback alongside the interface change.

A clean contract reduces the scope of code changes. It does not remove the lifecycle around the dependency or the need to prove that the replacement behaves correctly in production.

Principle 3

Optimize for the Next Engineer

Back to principles

Make the system understandable to someone who did not design it. Consistent structure, visible dependencies, and documented decisions reduce the time engineers spend reconstructing intent before they can make a safe change.

In practice: A new engineer should be able to trace an order from its entry point through business rules, persistence, integrations, and failure recovery. Names and module boundaries should make ownership and side effects apparent.

Tradeoff: Consistency can preserve a pattern that is not locally ideal, and documentation takes time. Standardize frequently used paths, document costly decisions, and allow exceptions with a clear reason.

Explore the reasoning: Optimize for the Next Engineer

Make dependencies and ownership visible

Use the same adapter structure across integrations and keep business orchestration in a predictable place. A pattern that varies between controllers, database hooks, and services forces engineers to rediscover the architecture for every feature.

Constructor parameters or explicit factories can show that a component needs an order repository, payment capability, clock, and logger. This supports isolated tests without global module patching; a large dependency-injection framework is not required.

Boundaries should answer where validation happens, who owns the transaction, whether an operation can be retried, and where side effects begin. If the answers depend on tribal knowledge, make the contracts clearer.

Reduce the cost of understanding a change

Use names that expose business intent: applyEligiblePromotions conveys more than processRules, and reserveInventory reveals more than updateStock. Linting and formatting can remove routine review decisions, but they cannot replace design review.

Document why a boundary exists, what an idempotency key protects, which system is authoritative, and how to recover a failed workflow. Restating function names adds maintenance without answering those questions.

Keep files aligned with coherent responsibilities. One function per file can create navigation cost; unrelated workflows in a 2,000-line service create change risk. Local development, tests, and fixtures should make the important contracts traceable without access to every production vendor.

Principle 4

Prefer Practical Simplicity

Back to principles

Choose the simplest architecture the team can operate confidently while meeting the business need. Complexity must justify its delivery and operating costs, including the engineering attention it consumes after launch.

In practice: A modular monolith can give pricing, orders, and fulfillment distinct interfaces while sharing a deployment. Extract a service when measured scaling, isolation, or ownership needs justify a separate runtime.

Tradeoff: Shared deployments can couple releases, compete for resources, and widen failure impact. Measure those costs against the timeouts, partial failures, data ownership, and operational work introduced by a distributed system.

Explore the reasoning: Prefer Practical Simplicity

Separate capabilities before deployments

Internal boundaries do not require network boundaries. A modular monolith lets capabilities have clear ownership while calls remain ordinary function calls. Refactoring across modules and running the application locally stay comparatively straightforward.

Keep dependencies directional and prevent database access or vendor calls from spreading across modules. One deployable application still needs disciplined boundaries; an unstructured monolith will become expensive to change.

Count the continuing operational work

Every service, queue, database, and cache needs monitoring, upgrades, capacity decisions, security patches, runbooks, and incident experience. Infrastructure spending captures only part of that cost. Small teams also pay with attention that could support business capabilities or known reliability risks.

Fewer moving parts can make request tracing and recovery easier. Shared resources can also cause contention, longer builds, coupled releases, and broader outages. Measure the actual constraints rather than assuming either deployment model is inherently simpler at every scale.

Know when a service boundary earns its cost

Independent scaling, dangerous workloads, different availability requirements, regulatory boundaries, and genuinely independent teams can justify separate services. Those benefits must outweigh the new work around authentication, schema compatibility, duplicate delivery, tracing, secrets, deployments, and cross-service consistency.

Principle 5

Configuration Beats Code

Back to principles

Let the business control frequent, well-understood variations without requiring an engineer and a production deployment for every change. Configuration should create a bounded capability with clear safeguards and ownership.

In practice: Marketing selects approved promotion conditions, benefits, limits, and dates. The engine validates the configuration and owns evaluation order, stacking, rounding, and explanations of the result.

Tradeoff: Safe configuration needs permissions, preview, audit history, and rollback. Building those controls may cost more than several hard-coded rules. Use configuration when change frequency or business value justifies it; keep unfamiliar or rarely changing behavior in code.

Explore the reasoning: Configuration Beats Code

Define a limited vocabulary

Pricing can represent base prices, effective dates, segments, channels, and precedence as validated data when the variations are known. Promotions can use conditions such as product attributes, quantity, and subtotal, with approved actions such as fixed discounts, percentage discounts, free items, or shipping benefits.

The engine must explain why a rule applied or did not apply. Keep evaluation order, conflicting rules, stacking, and rounding deterministic rather than leaving them to incidental implementation behavior.

Treat configuration as production behavior

Reject malformed promotions before they affect carts. Record who approved a price change, validate effective dates, and provide a way to preview and reverse changes. Moving a business rule into a form does not reduce the consequences of getting it wrong.

If configuration grows arbitrary branching, loops, remote calls, and custom scripts, it has become a programming language. It then needs versioning, testing, debugging, security controls, and engineers to maintain it. Ordinary code may be easier to operate.

Give feature flags a lifecycle

Flags support progressive delivery, operational control, experiments, and kill switches. They need owners and expected removal dates. Do not use them as substitutes for authorization, permanent product configuration, or unresolved architecture.

Long-lived flags add alternate paths to understand and test. Their operational value should exceed that ongoing cost.

Principle 6

Design for Small Teams

Back to principles

Build a platform the engineers delivering business features can also operate. Shared foundations should remove repeated decisions and operational work without requiring a specialist team for every capability or release.

In practice: A shared HTTP client supplies timeouts, bounded retries, correlation IDs, and redaction. A deployment template supplies health checks and rollback. Engineers inherit useful defaults instead of solving these concerns in every integration.

Tradeoff: Shared components can become bottlenecks or collect incompatible requirements. Keep their responsibilities narrow, version changes carefully, and allow justified exceptions. Consolidate operational concerns without forcing unrelated business domains into one generic model.

Explore the reasoning: Design for Small Teams

Standardize recurring production decisions

Useful foundations encode decisions engineers should make consistently. A job runner can provide idempotency, leases, retry limits, and dead-letter handling. API validation, migrations, error classification, structured logs, and ownership also benefit from predictable conventions.

A standard works when the safe path is easy to use. A document nobody follows does not provide the same capability, and a shared abstraction that makes application debugging opaque may simply move the work elsewhere.

Automate work with a high cost of error

CI should verify formatting, types, tests, builds, migrations, and dependency policy as appropriate to the repository. Deployment should produce repeatable artifacts and support rollback. Separate deployment from release when the workflow requires it.

Script local setup. Capture routine operational actions in safe tools or runbooks instead of relying on remembered commands. Evaluate productivity through the complete path from understanding a problem to deploying a safe change and diagnosing its result.

Limit the platform the team must support

Managed infrastructure can remove undifferentiated operations, but evaluate lock-in around core business capabilities. Multiple deployment, logging, queueing, or integration patterns each create maintenance work. Support alternatives when there is a concrete need and an owner for their cost.

Principle 7

Reliability Is a Feature

Back to principles

Design failure behavior as part of the product. Protect the customer's primary operation, make incomplete work recoverable, and keep failures visible to the people responsible for resolving them. A successful request is only part of the contract.

In practice: Keep optional recommendations off checkout's critical path. If fulfillment submission fails after an order is created, retain enough state to identify and resume that work without creating a duplicate order.

Tradeoff: Fallbacks, caches, and queues add alternate paths, stale data, or delay. Match those costs to business impact, make the customer-visible behavior explicit, and test recovery before relying on it.

Explore the reasoning: Reliability Is a Feature

Decide what can degrade safely

Payment authorization may be mandatory for checkout. Recommendations, analytics events, and marketing updates may not be. Make that distinction explicit and ensure optional dependencies cannot take down the primary operation. Silently returning incorrect information is not a safe fallback.

Caching can reduce vendor load and isolate latency, but acceptable staleness differs across product descriptions, inventory, prices, loyalty balances, and payment status. Define ownership, freshness, invalidation, and behavior for a cache miss during an outage. A cache must not become an untracked system of record.

Bound retries and handle uncertain outcomes

External calls need timeouts. Retry transient failures only when the operation is safe to repeat, using bounded attempts, backoff, jitter, and a total time budget. Validation failures should not be retried. Retries at nested layers can multiply traffic and worsen an outage.

An indeterminate payment response requires reconciliation before another charge is attempted. Idempotency protects against duplicate operations, while durable workflow state lets the system identify unfinished work. Asynchronous delivery also needs duplicate handling, retry limits, and a way to inspect and replay failures.

Make recovery observable

Use correlation and business identifiers in logs without exposing sensitive data. Measure customer and business outcomes alongside resource health. Recovery procedures should identify which steps completed and which can safely resume.

Circuit breakers, queues, and bulkheads are useful when they address a real failure mode the team can operate. Fallbacks need tests; idempotency records and workflow state need retention and cleanup policies. An untested recovery path remains an assumption.

Principle 8

Architecture Should Compound

Back to principles

Judge architectural investments by whether they make later work easier. Useful boundaries, standards, and operational capabilities should reduce the cost of change as the platform grows, while remaining understandable to the engineers who inherit them.

In practice: A repository boundary contains part of a database migration. A deployment template gives the next service safer defaults. A clear capability interface lets a new channel reuse business behavior without copying vendor assumptions.

Tradeoff: An abstraction without a real second use can preserve speculation. If a framework needs its original author to explain every change, its maintenance cost may exceed the leverage it provides.

Explore the reasoning: Architecture Should Compound

Test the investment against real changes

Ask whether the platform can accommodate:

  • Migrations: Stable business models and translation boundaries should contain changes to data stores, hosting, and frameworks.

  • New engineers: Consistent structure and documented decisions should let people contribute without repeatedly seeking the original author's permission.

  • New vendors: Contracts should preserve business intent and failure semantics while retaining provider context needed for historical support and reconciliation.

  • Acquisitions: Capability boundaries and authoritative data ownership should make overlapping catalogs, customers, and financial processes possible to integrate incrementally.

  • New products: Checkout, pricing, payment, order, and fulfillment capabilities should support new channels without inheriting all the assumptions of the first set of screens.

These changes still require work. The investment pays off when fewer unrelated parts must change at the same time and the team can understand and verify the transition.

Measure the return and repair what does not help

Look for lower effort on subsequent changes, fewer defects at established boundaries, faster onboarding, and safer operations. The number of patterns, services, or diagrams does not establish that return.

If each feature needs more coordination, exceptions, and knowledge of unrelated code, identify the boundary causing the cost. Repair it, improve the shared path, and remove accidental complexity. A rewrite is rarely the first response, and future benefits should not excuse failing to deliver value today.

The practical standard

My Definition of Great Architecture

Great architecture makes the next necessary change less expensive. It preserves business capabilities as vendors and technologies change, helps new engineers contribute, and makes failure and recovery understandable. Its tradeoffs fit the team that must operate it. The result is a platform the business can keep evolving without repeated rewrites.

See the principles under real constraints

The useful test of a principle is how it shapes a migration, integration, reliability decision, or platform boundary—not how convincing it sounds in isolation.