Atik Khatri
Solutions Architect
AWS
The v0.4 technical reference architecture: federated runtime cells, deterministic action execution, hospitality systems of record, property-level authority, operational flows, and testable conformance evidence.
How to build an AI platform that still works at property 500, across portfolios where the systems differ, the credentials belong to somebody else, and every conversation costs money.
Includes free, reusable tools for each repeatable part of the workflow.
Define once
Validate and execute safely
Map systems and fallbacks
Different systems, shared rules
Hospitality groups do not run one stack. They run a different stack at every property, and the question that decides whether an AI platform survives is not which model it uses. It is whether a capability that works at one property still works at property five hundred without being rebuilt.
Three things decide whether any of this works. Who is allowed to say yes, because in most portfolios the party operating a property is not the party holding the credential to its systems. Where each capability belongs, because scoping one wrongly means paying either for central infrastructure nobody needed or for the same capability rebuilt five hundred times. And what has to be true before a platform can be trusted with a write, because a model that can change a reservation directly is a model that will.
It is written for owners, operators and brand teams as much as for architects. Nothing in the main sequence names a product or a protocol. Where a point has a technical consequence, it is stated and then set aside.
A working demonstration is not a platform. The question is whether it still works at property 500, where the stacks differ, the credentials belong to someone else, and every conversation costs money.
Scalable is given a testable meaning here rather than a rhetorical one: the seven gates in step 09, four of them specific to AI. A platform meets that test when it can produce a current figure for each of them across the estate rather than for one installation. That is the definition this guide is organized around, and whether it is the right one is the fifth choice under what this guide does not decide. It is also why the cost step ends in figures rather than a claim.
The systems and use cases shown here are illustrative, not a comprehensive catalogue. They demonstrate how to apply the architecture to a few real operating situations. Add the workflows and system combinations relevant to your own portfolio; neither these examples nor the checklist cover every hospitality application.
The AI Orchestration Framework for Hospitality explains how AI agents and hotel systems coordinate safely: what an agent may access or change, how actions are authorized and sequenced, and how failures and handoffs are handled.
This guide focuses on making those capabilities work reliably across a hotel portfolio – clarifying responsibilities, connecting different property systems, protecting access boundaries, and managing performance, resilience and cost. It applies the orchestration principles rather than redefining them.
The two guides describe complementary responsibilities, not a required hosting layout. Components may run centrally, regionally or at a property, provided the same controls apply.
Three kinds of operator are in scope. A management company running a mixed portfolio of several brands. A mid-market brand whose properties are uniform. A franchisor that dictates the stack its franchisees run.
These need one architecture read through three authority models, not three architectures. The same layers, the same interfaces, and different answers to a single question: who is allowed to say yes.
Several brands under one operator. Authority is split by flag, and the operator cannot assume the portfolio is one pool.
One brand, consistent properties. Authority is concentrated, which makes most decisions easier and a few of them invisible.
The stack is dictated, but the franchisee owns the business. Authority is shared over the same guest.
What separates them is not technology. It is who holds the contract with each system, who can grant and revoke access, who holds the guest relationship, and who pays. Those four boundaries get different answers in each segment. The architecture states the answers per segment rather than forking into three designs.
Every step in this guide asks you to decide something once and hold it across the estate. Choose your authority model first. Otherwise, you will choose it accidentally, one capability at a time, and the answers will not agree with each other.

The matrix below states fourteen boundaries per segment. Each cell says where the boundary sits, not how the capability is built. Challenge the reasoned cells first, since those are derived from how the operating models work rather than observed directly.
| Decision | Management company, mixed portfolio | Mid-market brand, uniform portfolio | Franchisor dictating the stack | Basis |
|---|---|---|---|---|
| Tenancy grain, what counts as one tenant | The flag or the owner entity inside the portfolio. The operator is a parent, not a tenant. | The brand. Properties are scopes beneath a single tenant. | The franchisee legal entity. The franchisor is a policy authority, not a tenant owner. | Reasoned |
| Credential grant path, who can authorize access to the system of record | Frequently the brand rather than the operator, and it varies per management agreement. | The brand, centrally. One grant path. | Usually the franchisee, even where the franchisor mandates the product. | Portfolio |
| Data ownership, who holds the guest relationship | Each flag's guest data sits under that brand's agreement, so the operator cannot treat the portfolio as one pool. | One holder across the estate. Separating properties is a permission question. | The franchisee holds it. Franchisor access to aggregated data needs a stated basis. | Reasoned |
| Model selection and routing | Central, constrained by each brand's processing and residency terms. | Central, with regional variation where compliance requires it. | Central standard, but the franchisee's right to decline or substitute has to be written down. | Reasoned |
| Autonomy level, what an agent may do without a person | A floor per flag, since brand standards differ. A property may tighten it, never loosen it. | Central and uniform across the estate. | The franchisor sets the floor. The franchisee may tighten it and staffs the consequences. | Reasoned |
| Evaluation and quality thresholds | Central method, thresholds per flag, because guest expectations differ by brand. | Central, one bar for the estate. | Central method. The franchisee sees its own results, the franchisor sees conformance. | Reasoned |
| Agent voice, tone and languages | Brand-owned per flag. One agent voice across flags is not available to the operator. | Central, with local language coverage per property. | Franchisor sets the voice, the franchisee adds local specifics. | Reasoned |
| Workflows, prompts, guards | Central per flag. Maintained once per brand, never once per property. | Central. Property variation by parameter only. | Central template with a declared franchisee-editable band. Uncontrolled local editing is the failure mode. | Reasoned |
| Property and unit facts | Property. | Property. | Property, the franchisee site. | Portfolio |
| Human oversight and escalation | Policy central. Routing and queues local, and escalation targets differ by flag standard. | Policy and routing central. Queues local. | The franchisor sets a floor. It does not staff the desk and cannot assume response times. | Reasoned |
| Observability and audit access | Two audiences. The operator needs a cross-flag view, and each brand may require its own export. | One central view. A property sees its own slice. | The franchisee owns its audit record. The franchisor sees conformance evidence, not raw guest conversation. | Reasoned |
| Cost allocation | Recharged to owner entities, so cost has to be attributable per property or it cannot be billed at all. | Central budget. Cost per property is an internal efficiency measure. | A fee the franchisee can refuse. Cost per property is a commercial gate, not an optimization. | Portfolio |
| Change control and rollout | Central release with per-flag approval. The slowest brand governs the schedule. | Central and phased. | Adoption is contractual or voluntary. Rollout is a program, not a deployment. | Portfolio |
| Offboarding and exit | Properties change operator regularly, so offboarding is routine and needs a designed path. | Rare, and reasonably handled as an exception. | Deflagging happens, and the franchisee leaves with its data. | Portfolio |
This illustrative system map gives starting points for checking authority in each segment. It is not a complete system inventory or a universal ownership rule. Product selection, day-to-day administration, contractual ownership and permission to grant access can belong to different parties. Verify each against the actual agreement and record the result in the control boundary matrix.
| System area | Management company, mixed portfolio | Mid-market brand, uniform portfolio | Franchisor setting the stack |
|---|---|---|---|
| Property management (PMS) | Often brand-mandated per flag. Operator administration may not include the access grant. | Often a centrally held brand contract. | May be selected from an approved list while the franchisee holds the contract. |
| Central reservations (CRS) | Often controlled by each brand, outside the operator's direct authority. | Typically a central brand service. | Often operated by the franchisor. |
| Revenue management (RMS) | May be an operator-wide choice, subject to each flag's data permissions. | Often centrally selected and operated. | May be a franchisee choice within an approved list. |
| CRM, loyalty and guest profiles | Brand loyalty and local guest relationships may have different controllers. Do not assume cross-flag pooling. | A shared profile may be possible within the brand's permitted purposes and access rules. | Franchisor loyalty and franchisee guest records may be separate. Confirm both authorities. |
| Booking engine, channel manager and connectivity | Check brand control of branded inventory separately from the operator's unbranded inventory. | Often centrally controlled. | Often franchisor-controlled distribution, with access defined by agreement. |
| Guest apps and messaging | A brand app may coexist with local messaging. Identify the owner of each channel and thread. | Often shared centrally, with property-specific information and routing. | A franchisor app may coexist with franchisee channels. Assign ownership explicitly. |
| In-room technology, locks and energy | The asset owner may control these systems separately from the operator or brand. | Authority may sit with the brand or asset owner. | Check the franchisee and asset owner's responsibilities. |
| Food and beverage / point of sale (POS) | The operator or a third-party outlet operator may hold the system and grant. | Central or property-level responsibility may vary by concept. | Often franchisee-controlled, subject to the outlet arrangement. |
| Finance, accounting and labour | Often managed across the operator's portfolio, with owner-entity boundaries. | Often centrally operated. | Often franchisee-controlled. Franchisor access needs an explicit basis. |
| Reputation and digital marketing | Brand websites and operator-managed local presence may have different owners. | Often centrally managed, with local responsibilities defined. | Separate franchisor brand channels from franchisee local presence. |
The first three areas draw on portfolio experience; the remaining rows are reasoned starting points to test locally. An asset owner or third-party outlet is another authority to consult, not an additional deployment model imposed by this guide.
A central platform reaches only the systems whose credentials that party holds, and those sets barely overlap between segments. In two of the three, part of the stack belongs to a party outside the platform's control. A management company cannot acquire ownership of a brand's system of record by design. So the architecture has to operate through boundaries it does not control, which is why step 05 exists.
The previous step settled where authority sits. This one answers a different question, and it has to be asked about every capability separately: does this belong at a property, centrally, or somewhere in between.
The test is whether the value comes from what happens at one property, or from what happens across several. That gives three answers.
Works identically whether you have one property or five hundred. Nothing about it depends on the rest of the portfolio.
Does not merely improve with a view across properties. Does not function without one.
The logic is shared. What a guest or a member of staff sees is adapted per property.
Property scoped is cheaper to build first. So without a forced decision, capabilities drift to property scoped by omission, including the ones whose value only exists once they are centralized. You then pay either for central infrastructure you did not need, or for the same capability built again at every property.
Then there is a second step, and it is the one most easily skipped.
Read the relevant row of the control boundary matrix for each segment you operate in, and record the result. A classification is not finished until you have done this.
The capability test asks whether a capability benefits from scale. The matrix asks whether your segment permits it. A capability can pass the first and fail the second, and only the second is a constraint.
These four examples show the scoping test in use. They are not a complete list of use cases, and a system category does not determine the placement of every capability within it.
| Example | Starting placement to assess | What changes the answer |
|---|---|---|
| Housekeeping and engineering tasks | Property scoped when the task depends on that property's rooms and staff. | A shared procedure can be maintained centrally without merging local tasks or room facts. |
| Revenue management | A shared central method, with the actual forecasting or optimization capability assessed separately. | A mixed-portfolio operator may not pool booking history across flags. Central placement is not permission to pool data, nor proof that every revenue use case needs cross-property data. |
| Guest messaging and guest apps | Shared logic with local presentation and routing. | In a mixed portfolio, voice, tone and languages may be brand-controlled rather than property-tuned. That changes the boundary and its owner. |
| CRM and loyalty | Depends on the segment, contracts and authority for the guest relationship. | A single brand may share profiles within its permissions. A mixed portfolio may need separation by flag; franchisor loyalty and franchisee guest records may need distinct treatment. There is no universal placement answer. |
Record the result in a checklist that carries three fields per capability, not one: the scope classification with a one-sentence justification, the role bindings with a declared fallback where a role is unfilled, and the tenancy and isolation grain.
Treat the placement of loyalty as a recommendation to validate against your operating model. Check who controls the guest relationship, consent and the answer given to the guest before assigning a central or local role. A shared capability does not by itself authorize sharing guest data across brands.
Use the Deployment intake sheet in the Capability checklist workbook ↓ for each capability and release: business purpose → data readiness → model checks → controlled workflow, then release/review. State the outcome and owner first; check the necessary data and model evidence before deciding how the workflow is enabled.
Reference the completed boundary matrix, property bindings and scaling-gate evidence by workbook link, sheet, record and version. Keep the underlying evidence in its existing home instead of entering it again. Revisit this intake during readiness review and after a material change.
For shared instructions and knowledge, also identify the team authorized to maintain them. Maintain one controlled copy at that team or brand level, with permitted local variations. Keep property, unit and reservation facts scoped to the current task. This complements the capability test; it does not authorize pooling guest data or dictate where components physically run.
Pick a capability you have already built. Ask which of the three answers it is, then read the matrix row for your segment. If nobody can say which answer was chosen, it was chosen by whoever built it first.
Each fact has one authoritative owner. These are roles, not products: at each property, bind each role to the system responsible for that fact. One authoritative home does not mean one stored copy or one read path. Other systems may use derived, cached or replicated copies for performance, composition or resilience, provided the authoritative source, provenance, freshness and reconciliation rules are explicit. Copies retain the same access controls and access records; changes still go through the controlled write path to the authoritative owner.
| The fact | Where it lives | What every other system does |
|---|---|---|
| Conventional room-type availability | Recommended: property management or central reservations; select one authoritative owner | Reads from the authority or a governed copy. A copy does not become a second authority. |
| Rate and price | Revenue management | Quotes it. Never sets it. |
| Reservation state | Property management | Follows it. A change made anywhere else is a change that did not happen. |
| Guest identity and consent | Recommended: guest relationship management | References it. Consent is not collected again per channel. |
| Room and unit status | Recommended: housekeeping and maintenance | Reports into it and reads back from it. |
| Charges and folio | Recommended: property management | Requests a charge. Does not post one. |
| The guest conversation | Whichever channel currently owns the thread | Waits to be handed it. |
Drift appears when a second system answers from a stale or untraceable copy as though it were authoritative. An agent can amplify this by answering quickly and confidently from whatever it holds. Governed copies are not the problem; losing track of their source, freshness or reconciliation rules is. A read copy does not replace the authoritative commitment checks described in Step 08.
The table is a starting point for assigning authority, not a universal allocation. Guest identity and consent, room and unit status, and charges and folio use recommended role assignments. Validate them against how your estate operates. The availability row covers conventional room-type inventory, for which property management or central reservations is the recommended authoritative owner. Other sellable resources, attributes, services, operational states or entitlements may have different authoritative domain sources. A composed offer can have its own owner without taking ownership of the underlying resource facts. In every case, define the fact precisely, assign authority to the role that controls changes to it, not merely the system that displays it, and record the choice.
One more rule that is easy to miss. The guest conversation has one owner at a time, and handovers between channels are explicit events rather than something that emerges from whichever system replied last.
Choose one authoritative guest-identity record and map verified channel identities to it. An existing CRM or identity service can fulfil that role; the AI layer does not need a separate identity graph. A matching name, email address or shared phone number is not proof that two contacts are the same person. Verify control of the identities being linked and preserve consent, permitted purpose, and tenant, brand and property boundaries. R26 ↓
Keep inquiries and lead history in the designated lead or CRM system. Define the justified workflow event that creates or updates a PMS profile – for instance, a qualified reservation workflow – and reuse an existing verified profile where appropriate. Do not create a new PMS profile for every message, and make repeated requests safe against duplicate creation.
Ask the vendor: How do you link verified guest identities across channels without creating duplicate PMS profiles? Which system owns the identity record, and what triggers creation or updating of a PMS profile?
Test a verified guest moving from one channel to another and then booking, alongside a shared-contact case that must not merge. Use Reliability checks 41–44 in the Seven scaling gates workbook ↓; reference the existing authority and binding records rather than entering them again.
Pick one fact and trace every system that displays it back to its authoritative owner. For each copy, show its source, freshness limit and reconciliation rule, and test what happens when it becomes stale. Confirm that updates go through the controlled write path and that no copy can independently commit a competing version of the fact.
Fragmentation is not removed by architecture. It is relocated. The differences between your systems get absorbed once, low in the stack, or they get handled again in every workflow you ever write.
The map of hospitality systems carries around thirty categories, and most of them are acquiring their own agent. Across several hundred properties with stacks that do not match, one question governs everything: can one definition of a piece of work run everywhere without being rebuilt.
Workflows are defined in roles, not products. A definition names a price source, not a named revenue system. Each property binds those roles to whatever it actually runs.
That gives you a test you can apply to any workflow document you already have.
A workflow definition that names a product is not an estate capability. It is a property integration, and it will be rewritten for the next stack.

Two consequences worth planning for. A property missing a role degrades to a declared fallback rather than failing, and that fallback is a binding recorded in the bindings table, not a branch of logic buried inside a workflow. And definitions and bindings are versioned separately, because a change to a definition reaches every property while a change to a binding touches one.
Distinguish an estate-wide workflow from an identical feature set at every property. A common workflow can use the capabilities and declared fallbacks recorded in each property's bindings. Richer features can be enabled progressively where the systems and readiness checks support them, without forking the common workflow or waiting for every property. Claim a feature is available estate-wide only when every property supports it; a fallback is not evidence that the richer feature is available.
Search your workflow definitions for vendor names. Every hit is a place the definition will fork the next time a property changes system.
Every step so far produces a definition that is supposed to hold across the estate. This is the step that makes that possible, and it is the one most often missing.
The binding band records which product fills each role at this property, and the fallback where none does. One line per property. Everything above it is written in roles. Everything below it is whatever the property actually runs.
Workflows above have to know what each property runs, so the definition forks the moment two properties differ. You maintain five hundred configurations.
A property missing a system declares a fallback. A property migrating changes one line and nothing above it moves. Onboarding becomes a finite task with a definition of done.
The operating view below shows where the band sits. The six bands each answer one question: is this defined once for the whole estate, or does it record what is different about one property?

Take one capability and draw it through every band at two properties with different stacks. If anything above the binding band changed, you have described a system rather than a platform.
A connection alone does not establish readiness. In the Capability checklist workbook, use Prerequisite checks to ask the responsible system or agent vendor about the fields, APIs and events needed by the selected workflow. Record applicability, evidence, who verified it, and the reduced service or fallback when something is missing.
The checks cover selected booking and guest-service needs and four agent-vendor obligations; they are illustrative, not an exhaustive procurement checklist. Add checks for other use cases. Single-unit addressability matters only when the selected workflow needs it and does not redefine tenancy. A governed content copy or a separate payment service can be a valid design, not automatically a defect. R23 ↓
Every other step in this guide is about structure. This one is about safety, and it is the most technical step here. It earns that because it is the boundary an operator is most often asked to take on trust, and it is the one that should be least negotiable.
Models may propose structured actions. Only deterministic, policy enforced executors may authorize and perform side effects against business systems.
In plainer terms. A model can say what it thinks should happen. Something else, which is ordinary code with ordinary rules, decides whether it happens and then does it. The model never holds a credential and never reaches a system of record directly.
The path below is what sits between a proposal and a booking actually changing.

This is a recommendation about behavior, not a requirement about location. Every path that serves data should record who accessed what and under whose authority, using the same fields. Each path should use the same code to judge how current the data is, rather than implement its own version of that decision.
A central copy, a regional copy, a copy at the property, or a single shared path are all acceptable. The operator chooses where it runs; nothing here requires a component at a property. An existing central middleware layer does not need to be redesigned to follow this recommendation. The controlled write path described above stays unchanged.
Approval is a checkpoint on an action. Escalation transfers responsibility for the current task or conversation. It does not by itself transfer domain decision or write authority, ownership of guest data, or the broader guest relationship. A traveler's agent may still represent the traveler while hotel staff take over the task. On escalation the agent packages a summary and hands task responsibility to a role, never to a named individual, because staff rotate and no agent can track who is on shift. Existing permissions and consent boundaries still apply. After a transfer the conversation does not automatically come back, and no outcome status can be assumed to flow back either.
It is two in the morning and the night audit is running. A guest disputes a charge, and the agent proposes removing it. Follow the six steps and watch where it stops.
The proposal is well formed, so it passes the first check. The charge exists and the folio is open, so it passes the second. Policy identifies it as a financial write. Then the executor asks the property what state it is in, and the answer is that the business day is closing.
What happens next is the part most designs get wrong, and there are three ways to handle it.

Protect guest-facing response times when a publish, push or sync takes longer. Use durable background work that can resume after interruption, retry without creating duplicate writes, and hold repeatedly failing jobs for investigation. Separate workers and queues are one way to achieve this; their physical location remains the operator's choice.
Carry the tenant identity and permitted property scope through every request, job and tool. The tenant follows the authority model in step 01. Keep the same execution identifier from proposal through validation and the authorized executor to the resulting write, recording the outcome, elapsed time and cost. Test that a failure stays within its intended property or workflow boundary.
Three failure tests, all of which happen in real deployments. An escalation that dead ends because the person it was assigned to is off shift. Any component that waits for a completion signal that no operational system actually sends. And a financial write attempted during the audit window, to see whether it is refused, queued, or simply allowed through.
By the time an estate runs more than one agent, and most already do, two of them will eventually want the same resource at the same moment. That is not a rare collision to be handled when it happens. It is a design question with an answer, and the answer belongs in policy rather than in whichever component notices first.
A lock decides who was faster. It does not decide who is authoritative. Ordinary concurrency control will stop two agents corrupting a record. It will not tell you which of them should have won, and it will happily let the less authoritative one succeed because it arrived a moment earlier.
So precedence is resolved in advance rather than detected afterwards. Where two agents are each individually permitted to act on the same resource, policy says whose action takes effect before either of them runs.
That has two practical consequences for the write path in the previous step, and both are small additions rather than new machinery.
One thing this guide deliberately does not define: the authoritative list of which agents are permitted, and on whose authority. Treat that as a shared directory the platform consumes rather than a second list it maintains, because two lists means neither is authoritative. The fields this guide expects to consume are identity, owner, purpose, capabilities, versions and scopes.
Per-agent attribution is worth building early on operational grounds alone. Every question that follows a bad outcome starts with which agent did this, under whose authority, at what version. An estate that cannot answer that has to answer it by reconstruction, and reconstruction after the fact is a much larger job than recording it at the time.
Name two agents in your estate that are both permitted to act on the same resource, then ask what happens when they both do. If the answer describes a lock, a retry or a race, you have concurrency control and no precedence rule. Then ask whether your logs can name which agent performed the last write to any given record.
Every step so far would read much the same for a bank or a logistics operator. This one would not. Hospitality operations have shapes that general purpose platforms handle badly, and obligations that cannot be bolted on at the end. Both have to be in the design from the start, which is why they are a step and not an appendix.
The sequences below are normal operations, not edge cases. A platform that treats them as exceptions will work in a demonstration and fail in a hotel.
| Sequence | What actually happens | Why it is harder than it looks |
|---|---|---|
| Rate cascade | A pricing decision moves from revenue management through central reservations out to every distribution channel, and rate parity commitments have to hold at each hop. | One decision becomes writes to several systems that must all land or all be undone. A partial success is not a partial success, it is a parity breach. |
| Guest arrival chain | Check-in sets off room assignment, housekeeping notification, minibar activation, entertainment personalization and a welcome message. | The steps have an order, each can fail on its own, and stopping halfway leaves a guest in a room the rest of the estate does not know is occupied. |
| Night audit boundary | A nightly run closes the business day, posts recurring charges, advances reservations and resets counters. | It is a change of state for the whole property rather than a job that happens to be running. Some actions have to be refused while it runs. |
| Group block lifecycle | Held rooms with pickup deadlines, rooming lists, attrition penalties, and a cutoff date that releases whatever is left back to general sale. | It runs for weeks or months and each decision carries a contractual consequence, so it cannot be modelled as a request and a reply. |
| External booking arrival | Bookings from third-party channels arrive when they arrive, then become a reservation, a property record, a loyalty accrual and a confirmation. | Nothing is triggered by a person waiting for an answer, and deliveries repeat, so the same booking must never become two. |
| Displacement economics | Deliberate overbooking, with a real cost at the moment a guest has to be relocated or turned away. | An availability answer is not only a number. It carries a risk of walking someone, and that risk has a price that the answer should carry with it. |
| Cross-property transfer | A guest moves between properties in the same portfolio, and profile, preferences, folio and loyalty context need to follow. | The context has to cross exactly the boundary that the data rules are there to keep closed. |
| The rate-shift dilemma | A guest accepts a quote, but the price, cancellation terms or room availability change before the booking is committed. | Preserve the accepted offer, verify it with the booking authority, and use supported hold or conditional-commit controls. If it cannot be honoured, stop and request approval for a replacement. Payment success is not booking confirmation. |
Keep an immutable, expiring quote record with its identifier, property, dates, occupancy, room type, rate plan, price breakdown and cancellation terms. Record which version the guest accepted. A quote is not an inventory hold. Check the accepted terms and availability with the authoritative booking system at commitment, using its supported offer or inventory hold, or conditional booking operation, to prevent another sale between the check and the write.
If those controls are unavailable, disclose the limitation and use a controlled pending or manual route rather than promising a guaranteed booking. Do not assume a single atomic transaction spans reservations and payments. Honour a still-valid guaranteed offer; a change to the public rate alone does not cancel that guarantee. If the accepted offer has expired or can no longer be honoured, stop, issue a replacement and obtain approval before committing changed terms. R24 ↓
Track payment authorization, capture and reservation confirmation separately. Choose their order for the supported payment method and booking contract. If only one side succeeds, reconcile the actual outcome and apply the defined recovery – for instance, release an unused authorization or arrange a refund for a captured payment when the booking failed. Assign an owner and escalation deadline. For an unknown outcome, check the authoritative records before retrying; use the same operation identifier so a retry cannot create a second booking or charge. R25 ↓
Record the applicable tests under Reliability checks 35–40 in the Seven scaling gates workbook ↓. They cover changed terms, expiry, the last-room race, valid guarantees, partial success and uncertain outcomes.
Rate cascade and the group block lifecycle both sit on the booking critical path, and neither has an agreed specification for how systems should exchange them. They are named here because a plan that assumes they are solved will discover otherwise late. If you are buying, ask how each is handled rather than whether it is supported.

These are legal, contractual or safety obligations rather than business preferences. The distinction has a practical consequence: a preference can be checked at the end, an obligation has to be built into the path so it cannot be bypassed by a component that did not know about it.
| Obligation | Where it has to be enforced |
|---|---|
| Rate parity. Direct rates keep an agreed relationship to third-party rates. | Rate change proposals are checked against the agreements in force before anything is written to a distribution channel. |
| Brand standards. Franchised properties conform to mandated technology standards, data sharing terms and service levels by tier. | Brand policy applies as an overlay, with credentials scoped so that applying a standard does not hand over access. |
| Night audit lock. Financial writes are prohibited while the audit runs, and prior-day changes afterwards need specific authority. | The component performing the write checks audit state first. It is not enough for the requester to check. |
| Payment isolation. No model, prompt or retrieval step may reach raw card data, and no card data comes to rest in a retrieval index, a cache, a conversation log or a prompt store. | A structural boundary, not a rule. Payment components expose only a token and an authorization status, and a payment action crosses as a token rather than as an instrument. Treat that as a claim to be tested rather than a property you have. An index or a cache built before the boundary was in place can still hold card data, and the boundary says nothing about what is already stored. |
| Guest data jurisdiction. Regional privacy rules, including erasure that has to reach every system. | Residency enforced where data is accessed, and deletion propagated with evidence that it happened. |
| Overbooking exposure. Anything influencing availability has to quantify displacement risk. | Proposals carry a risk score, and past a threshold a person decides rather than the system. |
| Franchise data control. The franchisee legal entity controls guest data for its properties, and may be a different organization from whoever runs the property day to day. | Isolation at the franchisee entity, with property-level scoping beneath it. Where a management company operates a franchised property it acts on the franchisee's behalf and does not become the controller. |
| Operating mode. Properties run in modes such as peak, convention or renovation, with different parameters. | Mode is declared and shifts thresholds, approvals and permissions together, rather than each being tuned by hand. |
Deciding where a capability belongs, deciding who holds each layer for your segment, and the franchise data control obligation above are the same question at three stages. The first asks where a capability should sit. The second answers it for your operating model. The third is how the platform holds that answer once it is made, so that the boundary survives contact with a system that would happily ignore it.
Take the two sequences above that have no agreed specification and ask how each is handled end to end. Then pick the night audit lock and ask which component checks the audit state. If the answer is that the requesting application checks before it asks, the obligation is a convention rather than a control, and a future component that does not know the convention will breach it.
Then test payment isolation rather than assuming it. Search the retrieval index, caches, conversation logs and prompt store for data that looks like card details. Treat each match as a potential finding to investigate, not automatic proof of a leak. If no search has been run, record a verification gap: the boundary tells you what is not supposed to arrive and says nothing about what arrived before it was in place. Test the places data comes to rest, not only the steps that touch it.
Two things go wrong when operators budget for this. They count properties, and they count the cost of connecting systems. Neither is the number that scales.
The cost of adding a property is set by its stack, not its size. Two operators with four hundred properties each face different work depending on how many distinct systems sit underneath. Doubling inside systems you already support barely moves the curve. Doubling by acquiring properties on four unfamiliar systems moves it a lot.
Connection work is one time and falls as stacks repeat. Model cost recurs every month and tracks guests, not buildings. A five hundred property estate does not carry five hundred times the connection work. It does carry roughly five hundred times the conversations.

A capability can be affordable at pilot and unaffordable at rollout with no change to its design. How much context you assemble per turn, how often you reach for the largest model, and what share of requests you answer without a model at all are unit economics at estate volume, not engineering detail.
So the measure is cost per successful outcome, not cost per property. And there is a related requirement that is easy to treat as a reporting nicety when it is actually architectural.
In a management company, platform cost is recharged to owner entities. In a franchise model it is a fee a franchisee can decline. An architecture that cannot produce a defensible per-property cost figure cannot be bought in two of the three segments. Note the bound on the claim: this tells you whether a capability paid off. It does not make it pay off.
A platform meets this step when it can produce a current figure for each of the gates below. The first four are specific to AI systems. There are no target numbers here on purpose, and the reason is under what this guide does not decide.
| Gate | Measures | Passes when |
|---|---|---|
| Cost per successful outcome | Model and infrastructure cost of one resolved request, booking or answered question. | The figure is known, and falls as volume grows. |
| Escalation load | Share of interactions reaching a person, and the staff hours that implies across the estate. | Required human capacity is known in advance. |
| Evaluation coverage | Whether a change in answer quality at one property is detectable. | Quality is measured per property or per segment, continuously. |
| Model change resilience | Effect of a provider changing or retiring a model beneath a live workflow. | A model swap can be tested against real workflows before it becomes mandatory. |
| Onboarding effort | Hours to bring one property live within an already-integrated stack profile. First-time adapter work for an unfamiliar stack is measured separately. | Within a known stack profile, the figure at property 500 matches the figure at property 10. |
| Rollout reach | Time for one capability to reach every property entitled to it. | No per-property action is required for a capability to arrive. |
| Offboarding effort | Work to remove one property, and what remains afterwards. | Removal is a routine task with a defined end state. |
Before widening a rollout, test tool contracts and validators, denied access across tenant and property boundaries, simulated integrations, and whether prompts select the intended tools. Then run controlled live evaluations on a small, authorized scope. These checks strengthen evaluation coverage; teams still need representative questions and a method for judging answer quality.
Keep cost per successful outcome as the main measure, with breakdowns by property, workflow, model and vendor. Test whether budget controls can slow or stop one workflow without interrupting the rest. Use the workbook's Reliability checks sheet to record local acceptance criteria, actual evidence, owners and follow-up actions alongside the seven gates.
For each workflow, name the business outcome signal and the person responsible for measuring it. If the outcome cannot yet be measured, record why, who will address the gap and when it will be reviewed. A declared measurement gap makes the limitation visible; it does not pass the evaluation gate or count as a successful outcome.
Test model portability with a different model family, not only a version upgrade. Run representative workflows and compare answer quality, tool use, safety, latency and cost. Recheck scope boundaries and controlled writes. API compatibility does not establish equivalent behavior; record the experiment under Reliability check 29.
Keep operational telemetry – errors, latency, execution records and cost – distinct from guest-insight analytics or reporting sold as a product. The latter needs its own stated purpose, owner and access rules. Having reliability logs does not automatically authorize a new use of guest conversations or create a duty to deliver analytics reports.
Ask for the seven current figures. Not targets, not projections, the figures as they stand today. A gate nobody can produce a number for is a gate you are not measuring.

| Event | What has to happen | Done when |
|---|---|---|
| Joins | Its systems are identified and mapped to the roles the workflows use, and the history that comes with it is scoped. | Every role is either bound or has a declared fallback, and what history is imported has been decided rather than assumed. |
| Bound | Credentials issued, scope set, property facts loaded, historical data imported off the serving path, fallbacks recorded. | The property appears in the estate view, the import has finished, and no manual steps are outstanding. |
| Live | Capabilities enabled by wave rather than all at once. | Each enabled capability has a named owner at the property. |
| Changes | The affected bindings are swapped, and the history held in the old system is migrated, kept reachable, or deliberately let go. | Nothing above the binding layer had to be touched, and the agent knows as much about a returning guest as it did the week before. |
| Leaves | Access removed, data exported, records handed to whoever takes the property on, and anything under a retention obligation moved to a controlled hold. | Nothing relating to the property remains accessible to the departing operator or reachable from any live serving path, and someone can demonstrate it. Records retained for legal, audit or tax reasons remain, held segregated, access-controlled and off the serving path, with the retention basis recorded. |
Properties change systems while they are live. A property management migration, a reflag, a revenue system swapped after a contract review. Where workflows are written in roles this is a re-binding. Where they are written against products, every mid-life change costs what the original onboarding cost.
Two of these events deserve more than a table row.
Retention obligations, audit trails, tax records and the export owed to whoever takes the property on all survive departure by design. So the testable claim is about access and reachability, not existence. Two things have to be demonstrable: the departing operator can reach nothing, and no live serving path can reach anything belonging to that property. Whatever is retained sits segregated, access controlled, off the serving path, with a recorded reason for being kept.
A property management migration carries forward what is still ahead of you: reservations on the books, rates, inventory. What stays behind is what already happened. Past stays, folio history, the record of how a complaint was put right. The new system starts thin and the old one keeps the memory.
For an agent this is a visible regression. A returning guest recognized in March is a stranger in May. Across an estate where these changes happen monthly, estate memory becomes a patchwork of whenever each property last swapped a system.
This is why giving every fact one home needs a second clause. One home holds a live fact. History is what a system used to hold, and the system that created it may no longer be in that property's stack at all.
Where a property's history lives after a system change is not settled in current practice, and there are three defensible answers: migrate it into the new system, keep the old system alive read only, or hold it in the platform so it outlives the system that created it. This guide assumes the third for continuity across system changes, not because one authoritative home requires platform storage. Any of the three can preserve clear authority, provenance, controlled access and retention. Platform-held history can be costly, and your organization may choose differently.
What is not in doubt is the cost of not choosing. Without a decision, keeping the old system alive becomes the answer by default, discovered at the first migration and then repeated at every property.
Start with shadow operation: calculate proposed changes without publishing them externally. Record the target environment, properties and workflows, local acceptance criteria, observation period and evidence before enabling the next wave. The operator sets the thresholds and duration for its own risks and operating model.
Keep human confirmation where policy requires it. Enable automatic publishing only within explicitly authorized property and workflow scopes, and demonstrate that disabling it is faster than enabling it. Rehearse rollback or corrective action through the controlled write path. Record the rollout decision and outstanding conditions in the Seven scaling gates workbook ↓; repeat affected checks when a model, tool or binding changes.
A stack profile is the combination of systems a property runs. Group rollout waves by those combinations as well as operating permissions, rather than assuming region, brand or property size predicts integration effort. Identify uncommon combinations – the long tail – early. For each, choose a supported binding, a declared fallback, or a separately planned system migration.
Preparing a property should not automatically require replacing its existing systems. If a capability depends on a replacement, record that separate project and dependency rather than counting the capability as rolled out. Measure first-time adapter work separately from repeat onboarding within a known profile.
Import history on a separately budgeted path with scheduling, throttling and per-property limits. Protect both the guest-serving capacity and the shared systems of record: moving the loader to another worker is not enough if both still exhaust the same API limit. Test that one property's import cannot consume another property's service capacity.
Keep capabilities that require historical records off until their required import is complete, or explicitly enable only a restricted fallback that does not rely on the missing history. Record the import evidence and these tests under Reliability checks 31–34 in the Seven scaling gates workbook.
Take a property that changed its property management system in the last two years and ask what an agent can see about a guest who stayed before the change. Then ask who decided that, and when.
This example describes a hotel group operating its own portfolio: a single brand, roughly two dozen properties, and no franchising. Its production system illustrates several of the architectural choices used in this guide.
The example focuses on capabilities and operating decisions rather than named products or vendors.
Put a layer between the agents and the systems underneath, turning many point to point connections into far fewer. Treated the same guest question being answered differently on web chat and on messaging as unacceptable rather than untidy. Named an owner for every knowledge source and made retrieval strictly read only. Routed between models on cost, latency and quality with fallbacks. Put per agent permissions, audit logging and cost allocation in one place. Ran a steering group with declared decision rights and risk tiers. And ran agents from four different vendors against the same systems, on one stack.
Fragmentation is absorbed once, at step 04. Every fact gets one home, at step 03. Cost per successful outcome and model change resilience are measured, at step 09. And who is allowed to say yes is a standing process rather than a one-off judgement, at step 01. One rule from it has been adopted outright: knowledge is read only, teams keep working where they already work, content is pulled rather than pushed, and corrections happen at the source.
One segment only. A single brand operating its own portfolio is the middle column, where almost everything is central. It never exercises the credential grant path, per flag data ownership, or systems the operator does not hold.
Two dozen properties, not five hundred. It absorbs variance across system types. Absorbing variance across property stacks is a different problem, and a single brand operator largely does not have it.
Four vendors' agents, on one stack, against the same systems, in production today. Running more than one vendor's agent is not an advanced scenario you can defer. It is where operators already are, which is why deciding precedence between agents and being able to say which agent did what are early problems rather than late ones.
So read it as evidence that the approach works, not as a reference architecture to copy. The columns this example leaves untouched are exactly the ones most operators need.
Two different exercises get confused with each other, and it is worth separating them before you start.
Take three or four scenarios you actually expect to face, walk each one through the whole picture, and see where it breaks. This is cheap and you can do it this week.
A full evidence exercise against a named platform. Slower, and it needs the platform in front of you. Do not start here.
Conformance is demonstrated, not declared. A claim in a document is not evidence. And measure before you set targets. What follows says what has to be measurable. It deliberately does not tell you what the number should be, because nobody has enough comparable data across estates to publish one honestly.
Each one is answerable with evidence rather than an opinion. An answer of "we do not measure that" is a legitimate result and more useful than a reassurance.
Questions nine and twelve assume there will be somewhere to look up which agents exist and what they are allowed to do. The terms this architecture needs from that directory are identity, owner, purpose, capabilities, versions and scopes. This guide treats those fields as the expected interface rather than as a separate registry maintained by the platform. If the directory slips or changes, precedence between agents and per-agent attribution both need another home.
There is no total, no pass mark and no percentage, deliberately. The questions are not equally weighted and weighting them would depend on your segment, which is the thing this framework says you have to decide for yourself. Anyone who turns this into a number has added an assumption that is not here.
Review the Deployment intake in the Capability checklist workbook ↓ alongside these questions. Follow its evidence references to the boundary decisions, bindings, applicable prerequisite checks and scaling tests. Confirm the record matches the release and scope under review, with measurement gaps, exceptions, owners and next review dates visible. Completing a form is not release approval.
For booking or cross-channel guest workflows, also review the applicable quote-to-booking and identity checks in the Seven scaling gates workbook ↓ (Reliability checks 35–44). Reuse the intake, authority and binding evidence references; record a reason for any check that does not apply.
Legal duties depend on the use case, jurisdiction and your role. Have your legal or privacy team confirm what applies. This guide is design guidance, not legal advice.
Test whether rules can change without rebuilding the workflows. Configurable controls help you adapt; they do not by themselves establish legal compliance.
This guide supplies an architecture method and the tests that go with it. It does not replace decisions that belong to your own organization, or the authoritative sources that govern a jurisdiction, a contract, a brand agreement or a technical environment. Where it states one answer and practice offers several, those choices are set out in what this guide does not decide.
Who may grant access to a property's systems is set by management agreements and brand standards, not by architecture. This guide states where the boundary sits and what follows from it. It cannot tell you which side of that boundary your own agreements put you on.
Define operational thresholds for your own use and segment. This guide names what to measure and deliberately does not prescribe universal numbers, because the evidence does not support them. The one exception is tenancy leakage, where the acceptable target is zero.
The method requires a change in answer quality at one property to be detectable. It does not supply a method for detecting it. No test-set construction, no distinction between a snapshot and a dialogue, no paired negative cases. Treat this as the largest gap in the guide and expect to build your own.
Deciding whether a conversation is entitled to data is not the same as deciding whether that data is trustworthy. This guide covers the first. Third party text reaching a model is admitted on entitlement alone, and nothing here fences what it can then influence.
This guide explains the main architecture decisions and the checks to apply. It is not a detailed system specification. Your implementation still needs interface contracts, access policies, test cases and operating procedures suited to your systems. Use the linked workbooks to record decisions and evidence, and review the alternatives and limitations stated on this page before applying the recommendations.
It states how a platform is hosted, isolated, scaled, observed and tested. It consumes rather than redefines agent protocol behaviour, orchestration semantics, hospitality vocabulary and governance policy. Control implementation is out of scope.
The control boundary matrix, the capability checklist, the bindings table and the seven gates are supplied as accessible Excel workbooks beside the steps where each is used.
The method does not require a particular cloud, property management system, database, integration product or model provider. Product names, infrastructure choices and protocols are deliberately absent throughout, because they differ across the three kinds of operator this guide serves.
The official links in External resources are implementation examples, not mandatory workgroup requirements, a conformance profile or product endorsements. Select and pin only the versions that fit the deployment.
Three places where this gets genuinely difficult. They are stated plainly because a guide that only describes the tidy path is not useful.
The credential grant path determines whether a central platform is possible in the first place. A group running mixed flags often cannot authorize access to a property's system of record from the centre. The brand holds that right, and how much of it passes to the operator varies with each management agreement. Test this before the architecture hardens, because no amount of design gets around a right you do not hold.
Not because attribution is interesting, but because in two of the three segments the platform is recharged or is a fee someone can decline. If you cannot produce a defensible per-property figure, you cannot sell it internally.
Management contracts change and properties get deflagged. In a mixed portfolio this is routine business rather than an edge case, and it lands on precisely the things that are hardest to unwind: delegated credentials, accumulated history, shared configuration, and records that now belong to someone else. A design that treats exit as an exception will fail in the segment most operators are in.
This guide states one answer where hospitality practice currently offers several. That is deliberate, because a guide that hedges every step cannot be followed. This section is where the hedging goes instead.
Everywhere the method is settled, the guide states it plainly. Everywhere it is not, the guide still states one answer, and this is where you find out which ones those were. Five of them change what you would actually build, and each is flagged again at the step it affects.
A property's own messaging identity, and the chain that records who authorized an agent to act in a brand's name. Both matter, both are stated here, and neither has an established owner in current practice. If you are building now, you will have to assign them yourself.
Solutions Architect
AWS
The v0.4 technical reference architecture: federated runtime cells, deterministic action execution, hospitality systems of record, property-level authority, operational flows, and testable conformance evidence.
Senior Solutions Architect
AWS
The scalability framing across regions and operating contexts; the high-level architecture effort with David; and the requirement to apply the three-segment overlay across platform capabilities.
Senior Vice President, Travel & Hospitality vertical
Sirma Travel and Hospitality
Review from the Middleware Orchestration Layer Blueprint workgroup: the distinction between an agent registry and a model registry, authority-based precedence rather than concurrency alone, and alignment on deterministic writes, deployment neutrality, and the night-audit boundary.
Solutions Architect
AWS
Risk-driven sequencing and the initial five principles; editing and consolidation of the technical record; authorship of the single-page implementation guide; and clarification that consistent access records and freshness checks do not dictate where the system runs.
With Atik Khatri and Bharat Lakhiyani: The hospitality-specific architecture depth, including a taxonomy of the core operating systems and how AI workloads reach each; the sequencing that puts principles first; the requirement that the segment overlay applies to every capability; and the consolidation of every contribution into one document with its figures.
CTO
Polydom.ai
The architecture principles set, written as statements that can be tested; the data prerequisites and agent-vendor checks, including reduced service when prerequisites are missing; outcome measurement gaps, shared maintenance ownership, the telemetry/analytics distinction and model-family substitution tests; the boundary note with the orchestration work; and the four-pass editorial check: numbers recountable, cross-references correct, contribution claims confirmed with the person named, and unconfirmed positions labeled as proposals.
Program Manager
Sirma Travel and Hospitality
The control boundary matrix and its overlay on the landscape map, showing where the line between central and property sits per segment; the operating view this guide uses as its central figure; the seven scaling gates, four of them specific to AI, giving the word scalable a testable meaning; rollout by stack profile, import-capacity safeguards and the property-lifecycle figure; and the evidence-basis convention separating portfolio evidence from reasoning.
Founder
AI Hospitality Alliance
Workgroup direction and synthesis; the hospitality technology landscape map the control boundary work overlays, against which three segments are described from one system inventory; and the publication structure for the guide, including this format.
Co-Founder; workgroup facilitator
DominateAI
The capability scoping test, deciding whether a capability belongs at a property or centrally by where its value comes from rather than where its data sits; illustrative classifications for housekeeping, revenue management, messaging and CRM/loyalty, with the last dependent on segment and authority; and the position that placement follows accountability for the answer given to the guest, reaching the same conclusion by an independent route.
Vice President of Product Management
Milestone
Review input on authoritative sources, conflict resolution and data freshness; identity, access and allowed actions for external agents, including guest assistants; consistent hospitality entities and relationships; and separating operating models from platform choices. Proposed quality measures covering rates and availability, policies, amenities and services, appropriate escalation, and successful task completion.
CTO
ampliphi
The implementation-within-existing-systems framing, including model routing, evaluation and observability; and review input that vendor material must remain a labeled reference rather than a vendor endorsement. Pilot-to-portfolio reliability checks for durable execution, scoped access, evaluation, cost control and reversible rollout.
CEO & Co-founder
Lia
Review of booking and guest-identity failure points, with recommendations for handling quote expiry, rate and availability changes, separate payment and reservation states, and avoiding duplicate PMS profiles across channels.
President
GAIPAN
Reviewed the architecture's ability to support personalized hotel shopping, with recommendations on authoritative data ownership, governed copies, differing property capabilities, and separating responsibility for a task from decision authority and the broader guest relationship.
Founder
Punch Hospitality
Review of the guide's audience and accessibility, recommending a simpler executive view focused on business benefits and hotel examples while retaining technical depth and avoiding universal prescriptions.
Official sources and implementation examples for the boundaries discussed in the guide.
R1–R3 identify legal or payment-security boundaries that may apply. R4–R22 are examples of standards, specifications and platform foundations that an implementer may evaluate; they are not mandatory workgroup requirements, a conformance profile, or product endorsements. R23 is attributed implementation material; R24–R26 are provider documentation supporting specific design cautions, not required products. Select and pin versions for the actual deployment.
Reference for Regulatory note: Article 50 addresses disclosure and generated-output marking, with exceptions; Articles 12 and 14 address record-keeping and human oversight for high-risk systems. Human takeover is not a general Article 50 transparency duty. Applicability requires a deployment-specific legal assessment.
Reference for Regulatory note and property lifecycle: Article 5 includes storage limitation; rights, erasure obligations and international-transfer rules depend on the applicable circumstances. The guide's configurable-retention approach is an architectural interpretation, not statutory wording.
Reference for payment isolation under hospitality-specific flows. Consult the applicable PCI DSS version and assessment requirements with the responsible security team. Keeping raw card data out of model and retrieval paths is the guide's architectural boundary, not a claim of PCI compliance.
Risk-management foundation for governance, evaluation and evidence practices.
Management-system foundation for organizations providing or using AI systems.
Contract foundation for request-and-response APIs.
Contract foundation for event-driven interfaces.
Common event-envelope foundation.
Observability and telemetry vocabulary; implementations should pin the conventions they use.
Model-serving data-plane contract.
Workload-identity foundation.
Packaging and runtime portability foundation.
Software-supply-chain integrity and provenance foundation.
Hospitality and travel interoperability foundation.
Hospitality technology specification family.
Hospitality integration platform and API approach; inclusion is not a requirement or endorsement.
Example XML-based airline distribution standard that may inform distribution-adapter work.
Structured electronic-data-interchange foundation.
Publish-and-subscribe messaging protocol for connected-device and telemetry use cases.
Building-automation and control-network protocol.
Industrial-device protocol family; safety interlocks and command validation remain deployment responsibilities.
Privacy-information management foundation; it does not itself establish legal compliance.
Adapted for prerequisite review and the Capability checklist workbook under CC BY 4.0. Prompts were rewritten, applicability and evidence-reference fields added, and unsupported incidence claims and grading omitted. The adaptation does not imply endorsement or committee approval. This contribution is not an industry standard or a complete list of prerequisites.
Supports the booking and payment coordination distinction: distributed steps can need compensating actions, idempotent participants and concurrency controls. A saga does not itself provide transaction isolation or one atomic booking-and-payment commit.
Supports distinguishing authorization from capture in booking recovery. Holds depend on payment-method support and expiry. An authorization is not a reservation confirmation, and payment-provider capabilities must be checked for the actual integration.
Supports verification before linking identities: authenticate both accounts rather than treating a matching email as sufficient proof. This is a security implementation reference, not a requirement for an AI-owned guest database.