Menu

Software Architecture and System Design News

Latest curated articles from top engineering blogs

NetflixUberMetaLinkedInSpotifyGitHubAirbnbPinterestSlackDropboxCloudflareStripeDatadogFigmaShopifyAWSGoogle CloudAzureWerner Vogels& 15+ more

2428 articles

The New Stack·1h ago

Microsoft's Decision-1 Model: An In-house AI for Real-time Decision Making and Model Routing

Microsoft has developed Decision-1, an AI model for rapid, structured decision-making, offering an alternative to OpenAI's Decisions API. This model, post-trained on Alibaba's Qwen3.5-9B, aims to optimize internal operations like incident response and Copilot's on-device/cloud task routing, emphasizing cost-efficiency and latency improvements over larger LLMs. Its architecture and deployment strategy highlight a trend towards specialized, in-house AI components for critical system functions.

AI & ML InfrastructureDistributed Systems
15851
InfoQ Architecture·1h ago

Enhanced Event Buses in AWS EventBridge for Distributed Systems

AWS EventBridge has introduced enhanced custom event buses, significantly improving event-driven architectures in multi-account environments. These enhancements address complexities like managing subscriptions and cross-account routing, offering features such as ordered delivery, event replay, and content-based deduplication. This update simplifies building resilient distributed systems by providing a more robust and cost-effective event routing solution.

Cloud & InfrastructureDistributed Systems
171082
AWS Architecture Blog·1h ago

Self-Hosting AWS Regional Availability Data for Multi-Region Architectures

This article introduces open-source solutions for deploying AWS Regional availability data within a customer's VPC, enhancing governance and control for multi-Region architectures. It addresses the challenges of tracking service availability across AWS Regions for resilience, compliance, and expansion planning. The solutions provide a self-hosted dashboard and workload analysis tools to personalize availability insights.

Cloud & InfrastructureDistributed Systems
15908
Medium #system-design·1h ago

Designing a Production-Ready AI Support Agent

This article outlines the architectural considerations for building a production-grade AI support agent, emphasizing speed, safety, and cost-effectiveness. It discusses critical design choices related to LLM interaction, data retrieval, caching, and security, providing a practical blueprint for deploying AI solutions in real-world scenarios.

AI & ML InfrastructureDistributed Systems
16872
Cloudflare Blog·1h ago

On-demand CPU and Memory Profiling for Cloudflare Workers

Cloudflare introduces on-demand CPU and memory profiling with flamegraphs for Workers and Durable Objects, enabling developers to analyze resource usage directly in production. This feature is crucial for identifying performance bottlenecks and memory leaks in serverless functions, offering deep insights beyond aggregate metrics. The article highlights how this tool aids in optimizing performance and fixing memory issues in a distributed edge environment.

Performance & ScalingDevOps & SRE
12562
Cloudflare Blog·13h ago

Cloudflare's Clef-omni: A Multimodal Decision Model Architecture

Cloudflare introduces Clef-omni, an open-weight multimodal decision model capable of processing audio, video, image, and text inputs in a single API call. This architecture streamlines workflows by eliminating the need for cascading pipelines of single-modality models and leverages a mixture-of-experts (MoE) foundation for efficient, schema-constrained scoring. The article also highlights performance optimizations for existing Clef models, including faster inference and reduced pricing.

AI & ML InfrastructureDistributed Systems
1046850
AWS Architecture Blog·13h ago

Architecting Highly Available Oracle Databases on AWS with EVS and FSx for ONTAP

This article outlines a robust architecture for deploying highly available Oracle databases on AWS using Amazon Elastic VMware Service (EVS) and Amazon FSx for NetApp ONTAP. It details how to achieve sub-millisecond storage latency, cross-region disaster recovery, and integrate with existing VMware operational workflows, focusing on storage, compute, and networking considerations. The design emphasizes a hybrid storage approach using vSAN for ephemeral data and FSx for ONTAP for persistent, replicated database volumes.

Databases & StorageCloud & Infrastructure
1197280
Dev.to #architecture·13h ago

Designing Observability for Autonomous Agents: Beyond Simple 'SKIP' Logs

This article introduces "Observation Mode" for autonomous agents, a system design pattern focused on making an agent's decision to *not* act transparent and auditable. It addresses the problem of ambiguous "SKIP" logs by classifying non-actions into specific categories and requiring the agent to make and grade falsifiable predictions about system stability, enhancing trust and accountability in AI-driven systems through improved internal state reporting.

AI & ML InfrastructureDistributed Systems
1087320
Dev.to #systemdesign·13h ago

Designing Reliable AI Systems: Beyond Raw Models

This article shifts the focus from raw AI model capabilities to the critical aspects of designing reliable, controllable, and cost-effective AI systems for production environments. It emphasizes integrating AI models as components within larger, deterministic architectures, addressing challenges like hallucinations, cost, and data privacy. Key architectural patterns discussed include agentic workflows, GraphRAG for enhanced information retrieval, and the adoption of local SLMs for privacy and efficiency.

AI & ML InfrastructureDistributed Systems
1066767
Medium #system-design·13h ago

Monolith vs. Microservices: When Simpler Architectures Prevail

This article discusses the critical architectural decision between monolithic and microservices architectures, emphasizing that microservices are not a universal solution. It highlights scenarios where a well-designed monolith can offer significant advantages in terms of simplicity, development speed, and operational overhead, urging engineers to revisit the default assumption of microservices.

MicroservicesDistributed Systems
977615
The New Stack·13h ago

Building Feedback Loops for Production AI Agents

This article discusses the architectural considerations and operational challenges of establishing effective feedback loops for AI agents in production. It emphasizes connecting distinct stages like observation, data curation, and evaluation to continuously improve AI model performance and ensure alignment between AI engineering and SRE teams. The CoreWeave Forge platform is presented as an integrated environment addressing these handoff and data flow issues.

AI & ML InfrastructureDevOps & SRE
966822
Medium #system-design·1d ago

Database Internals for Portfolio Recommendation Apps

This article delves into the system design of a portfolio recommendation application, specifically focusing on how a single database can serve multiple roles: operational store, data lake, and data warehouse. It explores the architectural implications and internal mechanisms required to achieve this multi-purpose database utilization.

Databases & StorageDistributed Systems
1639812