RT
Infrastructure WednesdayOct 7, 20268 min read

Reference Architecture Decisions: Start Smaller Than the Diagram

Reference architecture diagrams are useful maps, but they are not workload requirements. Learn how to extract the mechanism, define boundaries, and choose the smallest reliable design.

Răzvan Todică, Senior Full-Stack Engineer and Team Lead
Răzvan Todică

Lead software engineer and technical consultant working across React, Next.js, TypeScript, Node.js, product delivery, and team leadership.

  • reference architecture
  • cloud architecture decisions
  • aws architecture diagrams
  • postgresql row level security
  • agent governance
Editorial cover for Reference Architecture Decisions: Start Smaller Than the Diagram
Original editorial cover generated for this article.
On this page
  1. Reference architecture decisions should begin with the workload
  2. Keep the diagram’s mechanism, not its visual authority
  3. A smaller architecture still needs explicit data boundaries
  4. Agent platforms expose the cost of invisible responsibilities
  5. Failure modes belong beside the boxes
  6. Choose the smallest reliable reference architecture
  7. Official sources

A reference architecture can make a difficult system look reassuringly complete. The boxes are already arranged. The services have names. The connections imply a path from input to storage, processing, access, and monitoring.

That convenience is also the risk.

A diagram is not a workload description, an availability requirement, a recovery plan, or a cost model. It shows one possible way to combine services. AWS describes its reference architecture diagrams as examples for combining AWS services across technology domains and industry verticals—not as universal system designs (AWS documentation).

The practical decision is therefore narrower than “Which cloud architecture should we adopt?” It is this:

Which parts of a reference architecture solve a named requirement, and which parts are merely inherited complexity?

That question matters to platform engineers, engineering leaders, founders, and budget owners because every additional component creates another responsibility: configuration, access control, observability, failure handling, retention, support, and eventual replacement.

Reference architecture decisions should begin with the workload

Before selecting services, write down the workload in plain language. The supplied AWS catalogue spans very different cases, including high-performance computing, geospatial imagery, fleet analytics, file sharing, and other domain-specific systems. Those examples demonstrate why a diagram designed for one workload should not silently become the baseline for another (AWS documentation).

Record at least these assumptions:

  • Workload: What is being served, processed, searched, stored, or automated?
  • Region: Where must the system operate? The supplied sources do not provide a region-specific price or limit, so no cost conclusion can be made from them.
  • Availability: What interruption can the product tolerate, and what must recover first?
  • Storage: What data is retained, for how long, and in which system?
  • Egress: What information crosses network or service boundaries?
  • Retention: Which data must remain available, and which can be deleted?
  • Support: Who owns incidents, upgrades, access reviews, and recovery exercises?

These are assumptions, not facts about a provider. They are the inputs required before a price, limit, or architecture comparison becomes meaningful. Without them, a diagram can only support a qualitative conversation.

Keep the diagram’s mechanism, not its visual authority

A useful reference diagram exposes a mechanism. For example, the AWS catalogue includes an architecture for connected aircraft that combines IoT Greengrass, Amazon S3, Amazon Managed Service for Apache Flink, and Amazon SageMaker AI for flight-data collection, analytics, and predictive maintenance (AWS documentation). The important lesson is not that every system needs those services. It is that a workload may have distinct stages for collection, storage, processing, and modelling.

Extract that mechanism and test it against your own requirements:

  1. Where does data enter the system?
  2. Which component owns durable storage?
  3. Where does transformation or analysis happen?
  4. Which output is operationally important?
  5. What can be delayed, replayed, or discarded?
  6. Which boundary needs explicit access control?

This approach preserves the useful reasoning while discarding components that have no named responsibility in the target workload.

The same discipline applies to architecture for AI agents. AWS’s agent material identifies recurring operational concerns: identity, memory, observability, governance, agent ownership, reuse, and cost attribution. It also describes a central registry as a way to see what exists, who owns it, and who has access (AWS Blogs). Whether a team uses that particular managed platform is a separate decision. The supported insight is that an agent system needs operational controls beyond a successful demonstration.

A smaller architecture still needs explicit data boundaries

“Start smaller” must not mean “leave access control for later.” The Supabase SolidJS quickstart provides a concrete example: create a Postgres table, grant only the privileges a role needs, enable Row Level Security, and define a policy for the intended access pattern (Supabase documentation).

That example supports a reusable architecture decision:

Make each data access path explainable as a role, a privilege, and a policy.

The example grants the anon role read access to a specific table and creates a policy allowing public reads for that table. This is not evidence that public reads are appropriate for every application. It is evidence that access should be expressed deliberately rather than assumed from the existence of a database or API.

A practical review can ask:

  • Which role is making this request?
  • Which table, function, or data boundary is exposed?
  • Is the granted privilege narrower than the role’s broader capabilities?
  • Is Row Level Security enabled where it is needed?
  • Does the policy match the product requirement, or merely make the demo work?
  • If the Data API is used, which tables or functions are exposed?

The smallest reliable solution may still contain several controls. What it should not contain is an unexplained collection of services copied from a diagram.

Agent platforms expose the cost of invisible responsibilities

The AWS agent source frames several common production problems: governance blind spots, agent sprawl, unpredictable cost, lock-in concerns, and isolated coding agents (AWS Blogs). These are not reasons to adopt a particular platform automatically. They are prompts for an architecture review.

For each proposed agent capability, name the owner and the control:

ConcernDecision to document
IdentityWhich agent can access which resource?
GovernanceWhich actions are allowed, denied, reviewed, or traced?
OwnershipWhich team is responsible for the agent after launch?
RegistryWhere is the inventory of agents and their access recorded?
ObservabilityWhich behaviour, failures, and evaluations are visible?
Cost attributionCan usage be assigned to a team, project, or action?
Model choiceWhat would have to change if the model or framework changed?

This table is an architecture boundary, not a promise of provider behaviour or a cost estimate. The supplied material mentions cost tracking and model flexibility, but it does not provide prices, regions, units, quotas, or observed results. Those details must be gathered separately before making a budget claim.

Failure modes belong beside the boxes

A reference architecture is incomplete if it describes only the healthy path. For every important component, write one failure question:

  • If ingestion stops, what data is lost, delayed, or replayed?
  • If processing fails, where is the unprocessed input retained?
  • If the database is unavailable, which user action is blocked?
  • If a policy is wrong, how is unexpected access detected?
  • If an agent behaves incorrectly, what trace or evaluation makes the problem visible?
  • If ownership changes, where does the next operator find the system inventory?

The sources support the need for observability, governance, identity, and evaluation in agent systems, and for explicit privileges and Row Level Security in the Supabase example. They do not specify recovery time objectives, backup methods, alert thresholds, or regional failover designs. Those must remain explicit project assumptions rather than invented defaults.

A concise operational checklist is therefore:

  • Name the workload and its critical path.
  • Record region, availability, storage, egress, retention, and support assumptions.
  • Give every component one stated responsibility.
  • Define roles, privileges, and data policies.
  • Document what happens when each critical dependency fails.
  • Specify how failures, access changes, and agent behaviour are observed.
  • Assign an owner for operations and recovery.
  • Separate provider list prices, estimates, calculations, and observed results.
  • Remove any component whose purpose cannot be defended in one sentence.

Choose the smallest reliable reference architecture

Use a reference diagram when it helps the team ask better questions, identify service boundaries, or compare possible mechanisms. Do not use it as evidence that all shown services are necessary.

A reasonable decision path is:

  1. Describe the workload. Avoid starting with a product catalogue.
  2. Extract the mechanism. Separate ingestion, storage, processing, access, governance, and observation.
  3. Map requirements to components. Every retained component needs a named responsibility.
  4. Test the boundaries. Review roles, privileges, policies, ownership, and data movement.
  5. Review failure and recovery. Put the failure mode next to the healthy path.
  6. Measure only after assumptions are stated. A price or limit without provider, unit, region or plan, and source date is not a decision-ready number.
  7. Prefer replaceable choices where uncertainty is high. The AWS agent material explicitly discusses lock-in concerns and model or framework changes; the decision should record what is portable and what is not (AWS Blogs).

The goal is not the fewest boxes. It is the smallest system whose responsibilities, access boundaries, failure modes, and operating ownership are clear enough to review.

Reference architectures are valuable maps. They become expensive when treated as destinations. Start with the workload, keep the mechanism, and make every remaining component earn its place.

If you are turning a promising diagram into an operational plan, I can help with a focused technical audit, architecture review, or decision record that keeps product and operational constraints visible.

Official sources

· Updated