A Pattern Language for Production LLM Platforms: Governed Routing, Agent Orchestration, and AI-Native Delivery
Description
A production platform built on large language models makes two kinds of decision, and most of its trouble comes from writing both into one clause. An optimization decision improves an objective: lower latency, lower cost, higher quality, fewer tests run. A boundary decision fixes a constraint that may not be relaxed for any gain: a residency rule, a least-privilege scope, a human-review threshold. When the two share a clause, improving one silently erodes the other, which is why efficiency and accountability are so often reported as a trade.
This specification is built on one invariant: a boundary is a clause the optimizer may not cross, and everything else is optimization. The contribution is a cross-layer architectural method for separating non-negotiable constraints from adaptive optimization and binding both to reconstructable evidence, applied identically across model routing, agent orchestration and AI-native delivery. The seventeen patterns are instances of that method rather than the contribution itself.
Each pattern is specified in the classical pattern form and carries three architectural declarations: the boundary it fixes, the optimizer it frees, and the evidence proving the boundary held. Every boundary is assigned to one of five classes covering data, authority, decision, resource and process constraints. Section 3 states the derivation method by which candidates were admitted or rejected, and publishes the rejections alongside the admissions so that the criterion can be examined rather than trusted.
Three mechanisms make the language operate as a language rather than a list. A pattern relationship graph names which pattern supplies the artifact, evidence or authority another depends on, including the single cycle by which a workflow improves from its own structural record and the economic chain running the full height of the stack. A normative event identity, with rules for causal parentage, retries, provider boundaries and retention, turns the requirement that evidence be joinable into something an implementation can satisfy or fail. And per-pattern applicability conditions replace categorical requirements, so that a pattern governing a mechanism an institution does not operate is out of scope rather than a gap.
Conformance is self-declared and published as a profile carrying the environment, the applicable set, per-pattern status, an evidence date and documented gaps. It is not a certification scheme, and no conformity assessment body operates against it.
The contribution is architectural rather than empirical. Every pattern carries an evidence level, and no pattern reaches the highest level, because no implementation unconnected to the author has been evaluated. Nothing has been measured. The specification separates what would falsify the invariant from what would falsify an individual pattern and from what would falsify the composition and adoption sequence, poses six research questions, and records the absence of a real implementation profile as a known deficiency of version 1.0.
An appendix reconciles the pattern identifiers with the names used across the author's papers and companion book series, including the acronyms PEVG and PARA, so that the two bodies of work can be cited as one.
Version 1.1 names two constructs the specification already contained. The central proposition is named the Boundary Invariant, and the three architectural declarations required of every pattern are together named the BOE Declaration. Neither carries a trademark, both are offered for use with attribution under this document's licence, and neither changes any requirement: the proposition, its wording and its priority date are those of version 1.0. Section 11 gains the two-family naming convention and a precedence rule fixing which document governs where this specification and the Defensible AI Framework Registry describe the same relationship.
Table of contents (English)
1. Introduction 2. Conceptual Model (Normative) 3. Pattern Derivation (Normative) 4. The Pattern Catalogue (Normative) 5. Composition (Normative) 6. Applicability (Normative) 7. Adoption Sequence 8. Conformance (Normative) 9. Related Work 10. Limitations and Known Failure Modes 11. Relationship to Other Frameworks Appendix A. Cross-Corpus Naming Crosswalk (Normative) Appendix B. Publication Provenance 12. ReferencesTechnical info (English)
Version 1.1 of the specification. 36 pages as rendered, with every figure as vector artwork rather than a raster image. 14 numbered sections and appendices, of which 6 are normative. 5 source SVG figures accompany the record and are reusable under the same licence. 41 references, each verified against a primary source. The Markdown source of record is deposited alongside the PDF, so the text is machine-readable without extraction.Notes (English)
Files
CHANGELOG.md
Files
(599.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:b30f2517fc111c72c291f0c517fa2934
|
9.1 kB | Preview Download |
|
md5:4e7224dfccb45b9f286720135fb0aef0
|
4.7 kB | Download |
|
md5:f3c7b04059a92e99f75cc2e26fc04d44
|
5.9 kB | Download |
|
md5:99b77de50c9fb4c70ad7d85ea29fd76b
|
6.7 kB | Download |
|
md5:65a8524d0046b09f53cfb11fe5e5c7c9
|
7.7 kB | Download |
|
md5:c198415389e429a155d979f0a82771b8
|
126.0 kB | Preview Download |
|
md5:bc84e246edecd5925fcc6375487b27c3
|
413.6 kB | Preview Download |
|
md5:20818a6e64adbcc64799299522968397
|
10.9 kB | Download |
|
md5:0f2ca1bbbf1bbe065d623a6ecfcbbbdf
|
7.5 kB | Download |
|
md5:eda3ddbaa126ab31cc2e87367fcbafdf
|
7.6 kB | Preview Download |
Additional details
Additional titles
- Subtitle (English)
- Pattern Language Specification, Version 1.1
- Alternative title (English)
- Production LLM Platform Architecture Patterns
- Alternative title (English)
- LLM Platform Design Patterns: Routing, Agent Orchestration and AI-Native Delivery
Identifiers
Related works
- Is documented by
- Other: https://nabeelkhan.com/frameworks/pattern-language (URL)
- Is supplement to
- Book: 978-1-0678317-1-4 (ISBN)
- Book: 978-1-0678317-2-1 (ISBN)
- Book: 978-1-0678317-3-8 (ISBN)
- References
- Report: 10.5281/zenodo.22109836 (DOI)
References
- Alexander, C., Ishikawa, S., and Silverstein, M. A Pattern Language: Towns, Buildings, Construction. Oxford University Press, New York, 1977. ISBN 0-19-501919-9. With Max Jacobson, Ingrid Fiksdahl-King and Shlomo Angel; Center for Environmental Structure Series, Volume 2.
- Alexander, C. The Timeless Way of Building. Oxford University Press, New York, 1979. ISBN 0-19-502402-8. Center for Environmental Structure Series, Volume 1.
- Gamma, E., Helm, R., Johnson, R., and Vlissides, J. Design Patterns: Elements of Reusable Object-Oriented Software. Addison-Wesley, Reading, MA, 1994. ISBN 0-201-63361-2.
- Buschmann, F., Meunier, R., Rohnert, H., Sommerlad, P., and Stal, M. Pattern-Oriented Software Architecture, Volume 1: A System of Patterns. John Wiley and Sons, Chichester, 1996. ISBN 0-471-95869-7.
- Liu, Y., Lo, S. K., Lu, Q., Zhu, L., Zhao, D., Xu, X., Harrer, S., and Whittle, J. Agent Design Pattern Catalogue: A Collection of Architectural Patterns for Foundation Model based Agents. arXiv:2405.10467, May 2024. https://arxiv.org/abs/2405.10467
- Cai, Y., Li, R., Liang, P., Shahin, M., and Li, Z. Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale. arXiv:2511.08475, November 2025. https://arxiv.org/abs/2511.08475
- Vandeputte, F. Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems. In Proceedings of the 2025 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software (Onward! '25), Singapore, October 2025, pages 44-62. DOI: 10.1145/3759429.3762620. Preprint: arXiv:2508.15411.
- Gulli, A. Agentic Design Patterns: A Hands-On Guide to Building Intelligent Systems. Springer Nature Switzerland, 2025. ISBN 978-3-032-01401-6 (softcover); eBook ISBN 978-3-032-01402-3; DOI 10.1007/978-3-032-01402-3.
- Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., and Dennison, D. Hidden Technical Debt in Machine Learning Systems. In Advances in Neural Information Processing Systems 28 (NIPS 2015), pages 2503-2511.
- Kreuzberger, D., Kuhl, N., and Hirschl, S. Machine Learning Operations (MLOps): Overview, Definition, and Architecture. IEEE Access, vol. 11, pages 31866-31879, 2023. DOI: 10.1109/ACCESS.2023.3262138.
- Niemen, G. How We Use Golden Paths to Solve Fragmentation in Our Software Ecosystem. Spotify Engineering Blog, 17 August 2020. https://engineering.atspotify.com/2020/08/how-we-use-golden-paths-to-solve-fragmentation-in-our-software-ecosystem
- ISO/IEC 42001:2023, Information technology - Artificial intelligence - Management system. International Organization for Standardization and International Electrotechnical Commission, Geneva, 2023. Catalogue entry: https://www.iso.org/standard/42001
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. Gaithersburg, MD, January 2023. https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
- Jitkrittum, W., Narasimhan, H., Rawat, A. S., Juneja, J., Wang, C., Wang, Z., Go, A., Lee, C.-Y., Shenoy, P., Panigrahy, R., Menon, A. K., and Kumar, S. Universal Model Routing for Efficient LLM Inference. arXiv:2502.08773, February 2025.
- Jain, K., Parayil, A., Mallick, A., Choukse, E., Qin, X., Zhang, J., Goiri, I., Wang, R., Bansal, C., Ruhle, V., Kulkarni, A., Kofsky, S., and Rajmohan, S. Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing. arXiv:2408.13510, August 2024.
- Chuang, Y.-N., Yu, L., Wang, G., Zhang, L., Liu, Z., Cai, X., Sui, Y., Braverman, V., and Hu, X. Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization. arXiv:2502.04428, February 2025.
- Li, B., Jiang, Y., Gadepally, V., and Tiwari, D. LLM Inference Serving: Survey of Recent Advances and Opportunities. arXiv:2407.12391, July 2024.
- Huang, K., Wu, H., Shi, Z., Zou, H., Yu, M., and Shi, Q. AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving. arXiv:2503.05096, March 2025. Accepted at ACM SoCC 2025.
- Sadhukhan, R., Chen, J., Chen, Z., Tiwari, V., Lai, R., Shi, J., Yen, I. E.-H., May, A., Chen, T., and Chen, B. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding. arXiv:2408.11049, August 2024.
- Xia, H., Yang, Z., Dong, Q., Wang, P., Li, Y., Ge, T., Liu, T., Li, W., and Sui, Z. Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding. In Findings of the Association for Computational Linguistics: ACL 2024, pages 7655-7671. https://aclanthology.org/2024.findings-acl.456/
- Lin, Q., Ji, X., Zhai, S., Shen, Q., Zhang, Z., Fang, Y., and Gao, Y. Life-Cycle Routing Vulnerabilities of LLM Router. arXiv:2503.08704, March 2025.
- Li, Z., Zhang, H., Han, S., Liu, S., Xie, J., Zhang, Y., Choi, Y., Zou, J., and Lu, P. In-the-Flow Agentic System Optimization for Effective Planning and Tool Use. arXiv:2510.05592, October 2025. The framework is named AgentFlow.
- Arora, D., Sonwane, A., Wadhwa, N., Mehrotra, A., Utpala, S., Bairi, R., Kanade, A., and Natarajan, N. MASAI: Modular Architecture for Software-engineering AI Agents. arXiv:2406.11638, June 2024.
- Niu, B., Song, Y., Lian, K., Shen, Y., Yao, Y., Zhang, K., and Liu, T. Flow: Modularized Agentic Workflow Automation. International Conference on Learning Representations (ICLR), 2025. arXiv:2501.07834. Note: this work defines workflows as activity-on-vertex graphs, which is the published grounding for AP-2.
- Zhang, J., Xiang, J., Yu, Z., Teng, F., Chen, X., Chen, J., Zhuge, M., Cheng, X., Hong, S., Wang, J., Zheng, B., Liu, B., Luo, Y., and Wu, C. AFlow: Automating Agentic Workflow Generation. International Conference on Learning Representations (ICLR), 2025 (Oral). arXiv:2410.10762. https://openreview.net/forum?id=z5uVAKwmjf
- Kaidhapuram, S. R. Human-in-the-Loop (HITL) Orchestration for Agentic Use-Cases: A Practical Framework for Supervising Autonomous AI Agents in Production Environments. International Journal of Computer Techniques, vol. 12, no. 6, 2025. Note: this is a low-visibility venue and is cited alongside [27] rather than alone.
- Mosqueira-Rey, E., Hernandez-Pereira, E., Alonso-Rios, D., Bobes-Bascaran, J., and Fernandez-Leal, A. Human-in-the-loop machine learning: a state of the art. Artificial Intelligence Review, vol. 56, no. 4, pages 3005-3054, 2023. DOI: 10.1007/s10462-022-10246-w.
- Pitkar, H. Platform Engineering and Developer Experience: A Systematic Review of Concepts, Benefits and Future Directions. World Journal of Advanced Engineering Technology and Sciences, vol. 18, no. 2, pages 241-248, 2026. DOI: 10.30574/wjaets.2026.18.2.0112.
- Baqar, M., Naqvi, S., and Khanda, R. AI-Augmented CI/CD Pipelines: From Code Commit to Production with Autonomous Decisions. arXiv:2508.11867, August 2025.
- Joshi, S. A Review of Generative AI and DevOps Pipelines: CI/CD, Agentic Automation, MLOps Integration, and Large Language Models. SSRN Electronic Journal, 2025. SSRN identifier 5290005.
- Soni, A. A., Parikh, M., Dhenia, R. N. K., Soni, J. A., Jha, A. R., and Shah, S. M. Reinforcement Learning for Dynamic Workflow Optimization in CI/CD Pipelines. arXiv:2601.11647, January 2026.
- Nabil, R. L., Zhu, H.-N., and Rubio-Gonzalez, C. CI-Bench: A Framework for Evaluating Large Language Model Tools on CI Failures. In Proceedings of the IEEE/ACM 48th International Conference on Software Engineering: Companion (ICSE-Companion '26), Demonstrations, Rio de Janeiro, Brazil, April 2026.
- Ghaleb, T. A. When AI Agents Touch CI/CD Configurations: Frequency and Success. arXiv:2601.17413, January 2026.
- Song, J. The Second Half of Cloud Native: The Era of AI-Native Platform Engineering Has Arrived. 2025. https://jimmysong.io/blog/cloud-native-second-half-ai-native-platform-engineering/ Note: this is a practitioner source cited for the established perception-action framing that OP-5 extends, not for the four-faculty model itself.
- Khan, N. A. LLM Systems in Production: Cloud-Native Patterns for AI Engineers. The Full-Stack AI Engineering Series, Book 1. iSystematic Inc., 2026. ISBN 978-1-0678317-1-4 (paperback); 978-1-0678317-5-2 (hardcover). Patterns IP-1 to IP-5 are developed here.
- Khan, N. A. Prompt Systems and Agent Orchestration. The Full-Stack AI Engineering Series, Book 2. iSystematic Inc., 2026. ISBN 978-1-0678317-2-1 (paperback); 978-1-0678317-6-9 (hardcover). Patterns AP-1 to AP-5 are developed here; PEVG is Chapter 3.
- Khan, N. A. DevOps for AI-Native Platforms. The Full-Stack AI Engineering Series, Book 3. iSystematic Inc., 2026. ISBN 978-1-0678317-3-8 (paperback); 978-1-0678317-7-6 (hardcover). Patterns OP-1 to OP-6 are developed here; PARA is Chapter 8.
- Khan, N. A. A Pattern Language for Production LLM Platforms: Governed Routing, Agent Orchestration, and AI-Native Delivery in Regulated Environments. Working paper, 2026. Not submitted; not peer reviewed; no preprint identifier assigned at the date of this document. See Appendix B.
- Khan, N. A. The Boundary and the Optimizer: A Pattern Language for Production LLM Platforms. Condensed magazine draft, 2026. Unsubmitted at the date of this document; no publication record. See Appendix B.
- A Pattern Language for Production LLM Platforms. Reference page: https://nabeelkhan.com/pattern-language-production-llm-platforms
- The MESA Framework. Reference page: https://nabeelkhan.com/frameworks/mesa