[{"content":"Notes and writing on backend engineering, AWS serverless, and platform engineering. The ideas and opinions are mine; the drafting is done with generative AI assistance. See the AI disclosure.\n","date":null,"permalink":"https://blog.rickgwaterman.com/","section":"Blog (Generative AI Assisted)","summary":"","title":"Blog (Generative AI Assisted)"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/categories/","section":"Categories","summary":"","title":"Categories"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/endpoint-management/","section":"Tags","summary":"","title":"Endpoint-Management"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/enterprise-it/","section":"Tags","summary":"","title":"Enterprise-It"},{"content":" Generative AI assisted. The ideas and opinions here are my own; an AI model helped draft and edit the text. For technical organizations, MacBooks should be evaluated as a first-class enterprise endpoint for engineering roles: manageable through modern identity and MDM controls, supportable inside Microsoft-centered environments, and worth standardizing when the goal is secure, low-friction execution.\nThat is the real shift.\nThe question is no longer whether Macs can be managed at enterprise standards. They can. The real question is whether the operating model around them has been modernized: enrollment, baseline policy, software delivery, access control, remote support, and lifecycle management.\nFor CTOs, security leads, and endpoint teams, that is the relevant business and technical decision. Done well, Mac support does not create an exception path. It reduces friction for technical staff while keeping control where it belongs: identity, device posture, compliance, and access.\nExecutive summary #If you only change five things, change these:\nEnroll every company-owned Mac through Apple Business Manager and Automated Device Enrollment. Use MDM as the authoritative control plane instead of manual setup documentation and one-off technician workflows. Tie device trust, identity, and offboarding together through your existing access model. Provide self-service software for approved tools instead of routing normal developer needs through tickets. Be selective about monitoring and endpoint controls that add friction without materially improving risk reduction. That combination addresses most of the operational drag teams encounter when Macs are treated as exceptions instead of managed platforms.\nMacs are already enterprise-manageable #Apple’s enterprise model is mature and well-documented. The building blocks already exist:\nApple Business Manager Automated Device Enrollment Mobile Device Management Declarative Device Management Platform SSO Managed software updates FileVault If your IT organization can manage Windows devices at scale, it can manage Macs at scale.\nThe practical challenge is usually not capability. It is operating model maturity. Many endpoint processes were built around older assumptions:\nMacs were exceptions rather than part of the standard fleet setup was technician-driven rather than enrollment-driven software delivery was ticket-first rather than self-service controls accumulated faster than they were rationalized hardware standards were based on office productivity rather than engineering workloads Those assumptions can be modernized.\nThe right operating model for Mac support #The easiest Mac fleet to support is the one designed around predictable, low-drama operations.\n1. Automated enrollment should be mandatory #Every company-owned Mac should be assigned to the organization before it reaches the user and should enroll during setup.\nThat removes the need for:\nwiki-driven onboarding hand-run bootstrap steps post-login technician intervention inconsistent setup states across users The target state is simple:\nUser opens the Mac -\u0026gt; authenticates -\u0026gt; baseline policy applies -\u0026gt; approved apps become available.\nThat is the standard a modern endpoint program should aim for.\nHelpful references:\nApple Platform Deployment Guide Automated Device Enrollment 2. MDM should be the control plane, not an afterthought #A good MDM program handles the repeatable baseline:\nFileVault and recovery key escrow passcode and screen-lock posture Wi‑Fi, VPN, and certificate delivery inventory and compliance software deployment wipe, return, and retirement workflows That is what MDM is for.\nIt should not become a dumping ground for every possible control simply because the platform allows it. The best MDM environments are consistent, supportable, and quiet.\nHelpful references:\nIntro to MDM Declarative Device Management 3. Identity should remain the primary control boundary #For security leadership, the durable control model is identity plus device posture.\nThe questions that matter most are:\nwho is the user is the device enrolled is it compliant what can it access how quickly can access be revoked This is also where Macs fit cleanly into Microsoft-centered environments. If the organization already uses Azure / Microsoft Entra ID, Microsoft Intune, and Conditional Access, Macs can plug into the same identity, compliance, and access model rather than living in a separate support silo.\nIn practice, that means:\nEntra ID can remain the identity and access backbone Intune can manage macOS enrollment and compliance Conditional Access can gate access to company systems Microsoft 365 remains fully usable for day-to-day business operations Helpful references:\nMicrosoft Intune macOS enrollment guide Microsoft Entra Conditional Access overview Platform SSO for macOS 4. Self-service software should be the default for common needs #If developers need a ticket for every browser, IDE, CLI tool, VPN client, or approved utility, the support model is generating its own load.\nA self-service software catalog produces immediate operational benefits:\nfaster onboarding fewer tickets better standardization less shadow IT For technical teams, this is one of the highest-leverage changes endpoint teams can make.\nSecurity should be strong, selective, and supportable #Macs can meet serious enterprise security requirements. The key is to prioritize controls that are enforceable, understandable, and operationally useful.\nA strong baseline usually includes:\nFileVault supported OS versions only managed software update deadlines secure lock and password posture MDM enrollment as a requirement for company access tightly controlled local admin rights sufficient logging for incident response and compliance fast disablement of accounts and tokens during offboarding This is also fully compatible with a Microsoft-centered security environment. Macs can participate in identity-aware access, compliance policy, Microsoft 365 collaboration, and managed browser usage without becoming second-class endpoints.\nA useful principle for security leads:\nSecure the device, secure the access path, and be cautious about controls that create more friction than measurable protection.\nWhy surveillance-heavy controls often produce the wrong trade #Many organizations add heavy endpoint monitoring with good intentions. They are trying to improve compliance, reduce risk, or increase visibility.\nThe problem is that, on developer endpoints, this often produces a poor trade:\ndegraded performance shorter battery life broken development workflows more support tickets lower trust between IT, security, and engineering That does not mean “no controls.” It means being more precise about which controls materially change risk.\nThe more useful questions are:\nIs the device enrolled? Is it encrypted? Is it patched within policy? Is access tied to identity and compliance posture? Can access be revoked quickly if needed? If those answers are strong, the organization is already covering the controls that matter most.\nMicrosoft 365 support is strong enough to remove the old objection #One of the oldest objections to Macs in enterprise IT is that Microsoft-heavy organizations must remain Windows-first to preserve productivity.\nThat is no longer a strong argument.\nMac users have solid support for:\nOutlook Teams Word Excel PowerPoint OneDrive Edge Microsoft 365 web apps For most business and technical collaboration, that is more than sufficient. In many environments, it is excellent.\nHelpful references:\nMicrosoft 365 for Mac Deployment options for Office for Mac Cross-platform secrets management still matters #Endpoint strategy works better when credentials and secrets workflows are also platform-neutral.\nA cross-platform password manager such as Bitwarden can simplify mixed environments by giving users a consistent vault model across Windows, Android, macOS, and iPhone.\nOperationally, that helps with:\nonboarding offboarding password hygiene browser support reduced platform-specific workarounds Helpful references:\nBitwarden Business Bitwarden downloads and extensions Personal Apple IDs should be treated as a policy question, not a blocker #For many organizations, the presence or absence of personal Apple IDs on work Macs becomes disproportionately contentious.\nThis is best treated as a policy decision.\nA company-managed Mac is still a company-managed Mac. Business data, required controls, and offboarding rules remain non-negotiable.\nBut a blanket prohibition on personal Apple IDs is not always necessary for technical staff, particularly when the goal is to reduce friction and improve usability.\nIf they are permitted, the boundaries should be explicit:\ncompany data remains governed by company policy required controls stay enabled unsupported sync behaviors are documented offboarding means access revocation and device reset Assumptions worth revisiting #“Macs are hard to manage” #Unmanaged Macs are hard to manage. Properly enrolled Macs are routine.\n“Macs are weaker from a security standpoint” #That is outdated. Modern Mac management supports encryption, update control, identity integration, compliance posture, and remote wipe effectively.\n“Mac support requires too much specialization” #Only if the operating model is inconsistent. A standardized Mac fleet with automated enrollment and self-service software is often easier to support than a fragmented mixed environment.\n“Developer preference is the only reason to support Macs” #Developer preference matters, but the stronger argument is workflow alignment: terminals, containers, cloud tooling, Unix-like environments, mobile ecosystem interoperability, and cross-platform engineering.\nRecommended defaults: enable, restrict, avoid #Enable # automated enrollment MDM-required management state FileVault with escrowed recovery keys managed software updates identity integration and SSO self-service app distribution remote support with user consent clean wipe, reissue, and offboarding workflows Restrict carefully # local admin access unmanaged access to sensitive internal systems unsupported OS versions ad hoc security exceptions Avoid unless clearly justified # broad spyware-style monitoring TLS interception on developer endpoints by default multiple overlapping endpoint agents manual setup paths that bypass enrollment ticket-only software delivery for common approved tools Hardware standards should match engineering workloads #This is one of the easiest ways to either reduce or create support burden.\nDevelopers do not use endpoints like office productivity users. They routinely run browsers, terminals, IDEs, local services, containers, collaboration tools, and security tooling in parallel.\nBuying weak hardware does not save much in practice. It usually shifts cost into:\nslower execution more complaints and exceptions shorter useful lifecycle additional support load Standard developer tier # MacBook Pro M5 Pro or higher 32 GB RAM minimum 1 TB SSD minimum Senior / staff / lead tier # MacBook Pro M5 Max or higher-end equivalent 64 GB RAM recommended 2 TB SSD recommended The point is not luxury. It is fit-for-purpose standardization.\nOnboarding and offboarding should be boring #That is a sign of a mature endpoint program.\nGood onboarding #The device is pre-assigned, enrolls automatically, applies baseline settings, and gives the user fast access to approved apps and services.\nGood offboarding #Identity access is disabled centrally, tokens are revoked, the device is wiped or returned through a standard process, and the Mac can be reissued without special handling.\nIf those workflows are still fragile, the right fix is process modernization.\nCritical next steps for CTOs and security/IT leadership #First 30 days # standardize the enrollment path define the baseline Mac policy inventory endpoint tools for overlap and friction publish an approved software model Next 60 days # tie Mac posture to identity and access decisions simplify remote support tighten offboarding and redeployment workflows document acceptable use clearly Next 90 days # standardize engineering hardware tiers remove unnecessary manual setup steps measure onboarding time, exception counts, and ticket volume reduce controls that add noise without improving outcomes Final thought #For technical leadership, the case for Macs is no longer a consumer-device discussion. It is an endpoint modernization discussion.\nMac management is mature. The opportunity now is to make it simple, supportable, and aligned with how modern technical teams work.\nThat means:\nautomated enrollment MDM as a real control plane identity-centered access self-service for common software needs selective, supportable security controls hardware standards aligned to engineering work When those pieces are in place, Macs stop being special-case devices.\nThey become what they should be: another well-managed, low-friction part of the enterprise environment.\n","date":"27 June 2026","permalink":"https://blog.rickgwaterman.com/posts/macbooks-in-enterprise-it/","section":"Posts","summary":"For engineering roles, MacBooks are a first-class enterprise endpoint — manageable through modern identity and MDM controls, supportable inside Microsoft-centered environments, and worth standardizing for secure, low-friction execution.","title":"MacBooks in Enterprise IT: A Practical Modernization Playbook for Technical Leadership"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/macos/","section":"Tags","summary":"","title":"Macos"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/mdm/","section":"Tags","summary":"","title":"Mdm"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/microsoft-intune/","section":"Tags","summary":"","title":"Microsoft-Intune"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/posts/","section":"Posts","summary":"","title":"Posts"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/security/","section":"Tags","summary":"","title":"Security"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/","section":"Tags","summary":"","title":"Tags"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/categories/technology/","section":"Categories","summary":"","title":"Technology"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/categories/architecture/","section":"Categories","summary":"","title":"Architecture"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/aws/","section":"Tags","summary":"","title":"Aws"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/cdk/","section":"Tags","summary":"","title":"Cdk"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/cost/","section":"Tags","summary":"","title":"Cost"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/nextjs/","section":"Tags","summary":"","title":"Nextjs"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/platform-engineering/","section":"Tags","summary":"","title":"Platform-Engineering"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/serverless/","section":"Tags","summary":"","title":"Serverless"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/terraform/","section":"Tags","summary":"","title":"Terraform"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/vercel/","section":"Tags","summary":"","title":"Vercel"},{"content":" Generative AI assisted. The ideas and opinions here are my own; an AI model helped draft and edit the text. Let me start with the part that is true without qualification: Vercel is great for frontend teams.\nIf you are a team of frontend engineers shipping a Next.js application, and nobody on the team wants to learn IAM, Vercel is the correct choice. It will be correct for a long time. Everything below the fold of this post is about a different team.\nWhat Vercel gets right #The developer experience is the product. Push a commit to any branch and you get a preview deployment with a branch-specific URL and a commit-specific URL, posted to the pull request. Zero configuration for Next.js. Instant rollback. Custom environments. None of it requires a line of infrastructure code, and that matters enormously for a team whose job is the interface, not the platform.\nNext.js is first-party. Vercel is \u0026ldquo;made by the creators of Next.js\u0026rdquo;, and it shows. The Next.js deployment docs describe a concept of verified adapters—platforms that run the full Next.js test suite—and Vercel is among them. Performance features like Partial Prerendering served from the CDN edge, ISR propagating globally in under a second, and skew protection land on Vercel first.\nThe engineering underneath is real. Vercel runs 126 points of presence and 20 compute regions. It built a Rust function runtime with in-function concurrency to get streaming and many-to-one invocation out of a substrate that did not natively offer it, and it reports that Fluid compute cuts compute cost by up to 85% for idle-heavy workloads. That is not marketing; that is hard distributed-systems work, done well.\nIf that is the whole story for your team, stop reading and go ship.\nNow the other team #Here is the team this post is actually for: an organization with a platform or backend group that already writes AWS CDK or Terraform, already runs Lambda and CloudFront, already has an AWS account structure with IAM boundaries and CloudWatch dashboards and a WAF. For that team, Vercel\u0026rsquo;s premium buys much less than it looks like, and costs more than the invoice shows.\nYou are paying a markup on your own primitives #Vercel\u0026rsquo;s region table maps iad1 to us-east-1, pdx1 to us-west-2, fra1 to eu-central-1. All twenty regions are AWS regions. Vercel has said plainly that Vercel Functions run on AWS Lambda — \u0026ldquo;we\u0026rsquo;ve turned lambda into an edge-first compute layer\u0026rdquo; — reached through a TCP tunnel Vercel built between its infrastructure and Lambda.\nCompare the list prices, as of this writing.\nResource Vercel Pro AWS direct Function invocations $0.60 per million after 1M included $0.20 per million after 1M free Data transfer out $0.15 per GB after 1 TB included $0.085 per GB after 1 TB free Edge / CDN requests $2.00 per million after 10M included $1.00 per million after 10M free Three times the request price, roughly 1.8 times the egress price, twice the CDN request price. That is the cost of the renaming layer. If your team cannot write the CDK to put a Lambda behind CloudFront, the markup is a bargain. If it can, the markup is paying for something you already own.\nAnd then there are seats. Vercel Pro is $20 per developer per month. That number scales with headcount, not traffic. A platform team of twelve pays $240 a month before a single request is served. CloudFront, Lambda, Amplify Hosting, and an SST deployment have no seat price at all.\nThe controls you want are behind the Enterprise wall #Look at what the Enterprise plan gates: SSO/SAML login, directory sync, audit logs, log drains, tracing support, Datadog and New Relic integrations, SLAs, automatic failover regions, and Secure Compute. Pricing is \u0026ldquo;Custom.\u0026rdquo;\nSecure Compute is the one that matters most to the AWS team. It is how your Vercel functions reach a private RDS instance or an internal service. It is \u0026ldquo;available as an Enterprise feature,\u0026rdquo; it works by VPC peering from a Vercel-owned network into your AWS VPC, it does not support the Edge runtime or middleware, and private data transfer that leaves the network is billed at $0.15 per GB.\nOn AWS the Lambda is already in your account. Put it in the VPC. Attach the security group. Write the IAM policy scoped to the one queue it needs. Ship logs to the CloudWatch group your platform team already alerts on. Trace it with X-Ray. Every one of those is a few lines of CDK, none of them requires a sales call, and the data never leaves your account. SST\u0026rsquo;s migration guide says it in one sentence: \u0026ldquo;Your data will never leave your AWS account.\u0026rdquo;\nSpend controls lag, and you do not own the throttle #In June 2024 the art platform Cara grew from 40,000 to 650,000 users in a week, peaked at 56 million function invocations a day, and received a $96,280 bill from Vercel for serverless function execution (Silicon Republic). The founder\u0026rsquo;s own account: \u0026ldquo;I was hit with a $100k bill from one of our service providers for 1 week of traffic.\u0026rdquo;\nVercel responded with Spend Management, which is better than nothing and worse than you would hope. From the docs: \u0026ldquo;Setting a spend amount does not automatically stop usage.\u0026rdquo; Pausing is opt-in, pauses all production projects with a 503 DEPLOYMENT_PAUSED, and \u0026ldquo;because these checks are not continuous, notifications, webhooks, and project pausing can trigger several minutes after you cross your spend amount.\u0026rdquo; Seats and add-ons are excluded.\nUsage-based billing without a hard cap is a property of cloud generally, not Vercel specifically — Yan Cui made that point at the time, fairly. The difference is who owns the throttles. On AWS, the team that writes the CDK sets reserved concurrency on the function, puts a rate-based rule on the WAF in front of CloudFront, and can choose CloudFront\u0026rsquo;s flat-rate plans with no overage charges. Those knobs exist on Vercel only to the extent Vercel exposes them.\nThe bill is hard to model #Vercel has repriced its compute three times in about fifteen months: metered SKUs in April 2024, Fluid compute in February 2025 promising up to 85% savings, and Active CPU pricing in June 2025 promising up to 90% for idle-heavy workloads. Each was framed as a cut. Each changed the unit you are billed on — from GB-hours to Active CPU hours plus provisioned memory plus invocations plus edge requests plus origin transfer.\nLambda\u0026rsquo;s GB-second model has not changed in years. A platform team can forecast it on a napkin. That stability is a feature for anyone whose job includes a budget.\nSelf-hosting Next.js caught up #The strongest historical argument for Vercel was that Next.js outside Vercel was a degraded experience. That argument has expired.\nThe Next.js team\u0026rsquo;s own platform guide now says: \u0026ldquo;To run Next.js, your platform needs a Node.js server. That\u0026rsquo;s it.\u0026rdquo; A single next start \u0026ldquo;handles every Next.js feature correctly: Server Components, ISR, PPR, Cache Components, Server Actions, Proxy, and after().\u0026rdquo; And: \u0026ldquo;There are no private framework hooks or integration paths: Vercel\u0026rsquo;s adapter uses the same public API as every other adapter.\u0026rdquo;\nOpenNext, whose AWS adapter is maintained by the SST community, covers App and Pages Router, SSR, SSG, ISR, middleware, image optimization, and use cache. SST wraps the roughly seventy AWS resources that make up a Next.js deployment into one component: new sst.aws.Nextjs(\u0026quot;MyWeb\u0026quot;, { link: [bucket] }), deployed to Lambda, CloudFront, and S3 in your account. Amplify Hosting supports Next.js 15 \u0026ldquo;without the need for an adapter,\u0026rdquo; with pull-request previews and atomic deployments, at $0.30 per million SSR requests and no seat fee.\nIf you would rather own it yourself, the pieces are a CloudFront distribution, an S3 bucket for static assets, a Lambda function URL for the server, and a cache handler. A team with CDK fluency has built that before, for something else, and has the constructs lying around.\nHonest counterpoints #Engineering time is not free. SST\u0026rsquo;s \u0026ldquo;around seventy low-level AWS resources\u0026rdquo; is the honest count. Self-hosting means you own multi-instance cache coordination, the Server Actions encryption key, deploymentId for skew protection, and the fact that an ALB with Lambda integration may buffer streaming responses. These are solved problems, but they are your problems now.\nParity is functional, not always performance. The Next.js docs distinguish functional fidelity from performance fidelity. PPR\u0026rsquo;s static shell at CDN latency and sub-second ISR propagation are what Vercel tunes. OpenNext\u0026rsquo;s page still says \u0026ldquo;some features are work in progress,\u0026rdquo; and Amplify is \u0026ldquo;not verified by the Next.js team.\u0026rdquo;\nIdle-heavy workloads can genuinely be cheaper on Active CPU billing. If your functions spend most of their wall-clock waiting on a model API, paying only for active CPU is a real advantage over provisioned compute. Lambda bills wall-clock. For that specific shape, run the numbers both ways.\nThe decision #The question is not \u0026ldquo;is Vercel good.\u0026rdquo; It is. The question is whether your organization already has the specialization Vercel is charging you not to need.\nA frontend team without it: use Vercel. You will ship faster and the premium is cheap against the alternative of learning AWS under deadline.\nA team with CDK or Terraform depth: the markup buys a renaming layer over Lambda and CloudFront, the controls you care about are behind an Enterprise contract, the spend cap is a notification, and the framework now runs everywhere on the same public API. Put the Next.js app in your own account, next to the rest of your infrastructure, governed by the same IAM, observed by the same dashboards, and paid for at list price.\nYou already did the hard part. Do not pay someone else for it twice.\n","date":"24 February 2026","permalink":"https://blog.rickgwaterman.com/posts/vercel-is-great-for-frontend-teams/","section":"Posts","summary":"Vercel\u0026rsquo;s developer experience is the best in the business, and for a frontend team without cloud depth it is the right call. But if your organization already has AWS CDK or Terraform specialization, you are paying a markup on Lambda and CloudFront for an abstraction you do not need, with spend controls that lag, private networking behind an Enterprise contract, and a self-hosting story that has finally caught up.","title":"Vercel Is Great for Frontend Teams"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/career/","section":"Tags","summary":"","title":"Career"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/infrastructure-as-code/","section":"Tags","summary":"","title":"Infrastructure-as-Code"},{"content":" Generative AI assisted. The ideas and opinions here are my own; an AI model helped draft and edit the text. There is a common worry among engineers who go deep on AWS: that specialization is a trap, that the real skill is \u0026ldquo;the cloud\u0026rdquo; in the abstract, and that tying yourself to one provider makes you brittle.\nThe opposite is true, and the evidence is sitting in the documentation of every developer platform you have ever used. Open the regions page. Read the codes.\nThe region table is the tell #Vercel\u0026rsquo;s region docs list a \u0026ldquo;Region Name\u0026rdquo; column. The values are iad1, sfo1, pdx1, dub1, fra1, hnd1. The column next to it says us-east-1, us-west-1, us-west-2, eu-west-1, eu-central-1, ap-northeast-1. Twenty compute regions, every one of them an AWS region code with a friendlier alias. Vercel said it directly in a 2022 post written with AWS: \u0026ldquo;We\u0026rsquo;ve turned lambda into an edge-first compute layer,\u0026rdquo; running \u0026ldquo;millions of functions that get invoked over five billion times per week.\u0026rdquo;\nNetlify\u0026rsquo;s function configuration lists deployable regions as cmh | US East (Ohio), iad | US East (N. Virginia), pdx | US West (Oregon), dub | EU (Ireland). Those are not Netlify\u0026rsquo;s names for places. They are AWS\u0026rsquo;s display names, verbatim. The same page says your Node version \u0026ldquo;must be a valid AWS Lambda runtime for Node.js,\u0026rdquo; set through an environment variable called AWS_LAMBDA_JS_RUNTIME.\nSupabase: \u0026ldquo;Supabase will deploy your project to an available AWS region within that area based on current infrastructure capacity.\u0026rdquo; The region list is us-east-1, eu-central-1, ap-southeast-2, and so on.\nHeroku\u0026rsquo;s region docs show the Platform API returning \u0026quot;provider\u0026quot;:{\u0026quot;name\u0026quot;:\u0026quot;amazon-web-services\u0026quot;,\u0026quot;region\u0026quot;:\u0026quot;eu-central-1\u0026quot;}. Render: \u0026ldquo;Render runs in the very same data centers as your EC2 instances, Lambda invocations, and other AWS resources.\u0026rdquo; PlanetScale prefixes its region slugs with the provider: AWS ap-northeast-1, gcp-us-east4.\nMove up the stack to the data layer and the pattern holds. Snowflake \u0026ldquo;runs completely on cloud infrastructure\u0026rdquo; and is hosted on AWS, Azure, or Google Cloud. MongoDB Atlas and Confluent Cloud list the same three and use the providers\u0026rsquo; own region codes in their tables.\nNone of these companies are hiding it. The abstraction is a renaming layer and a very good developer experience on top of primitives you can rent directly.\nPlatform limits are AWS limits in a costume #Once you know the substrate, the platform\u0026rsquo;s ceilings stop being surprising.\nNetlify\u0026rsquo;s Lambda compatibility page states that the total size of all environment variables \u0026ldquo;cannot exceed 4 KB\u0026rdquo; because that is \u0026ldquo;AWS Lambda\u0026rsquo;s environment variable size limit.\u0026rdquo; Vercel\u0026rsquo;s Fluid compute, launched last month, exists so that \u0026ldquo;multiple invocations can share the same physical instance\u0026rdquo; — which is to say, it exists to escape Lambda\u0026rsquo;s one-invocation-per-execution-environment model, the model every AWS engineer has been designing around for a decade.\nIf you know Lambda\u0026rsquo;s cold-start behavior, its payload limits, its execution-environment lifecycle, and its regional quotas, you already know where Vercel Functions and Netlify Functions will bend. You do not need the vendor\u0026rsquo;s documentation to tell you; you need it to tell you which of the limits they have papered over and how.\nOutages propagate downward #When the substrate fails, the platform can only wait. Heroku\u0026rsquo;s 2017 post-incident review is the most honest version of this I have read: \u0026ldquo;Many of the problems during this incident stemmed from the Amazon S3 outage on the 28th,\u0026rdquo; and then: \u0026ldquo;Any instability or unavailability due to issues with those providers or technologies is a consequence of our choices.\u0026rdquo; The AWS summary of that event explains what actually happened: a command input \u0026ldquo;entered incorrectly,\u0026rdquo; and \u0026ldquo;a larger set of servers was removed than intended.\u0026rdquo;\nIn December 2021 Heroku\u0026rsquo;s status page reported that \u0026ldquo;our upstream provider is experiencing elevated error rates in the US region,\u0026rdquo; affecting apps \u0026ldquo;in both the US and EU regions.\u0026rdquo; Heroku never named the provider. The AWS post-event summary for the same afternoon describes an automated scaling activity that \u0026ldquo;triggered an unexpected behavior from a large number of clients inside the internal network,\u0026rdquo; with EC2 API errors starting at 7:33 AM PST.\nIn June 2023 the Lambda frontend fleet in us-east-1 \u0026ldquo;crossed a capacity threshold that had previously never been reached,\u0026rdquo; triggering \u0026ldquo;a latent software defect.\u0026rdquo; EventBridge delivery latency reached 801 seconds. That is the service Vercel and Netlify functions run on, in the region Vercel functions default to.\nAn engineer who can read aws.amazon.com/message/* understands their own incident before their vendor\u0026rsquo;s status page finishes updating. That is not a small thing during an outage.\nUndifferentiated heavy lifting is a dial #Werner Vogels\u0026rsquo; framing in the 2014 Lambda announcement — customers want to focus on \u0026ldquo;their unique application logic and business needs, not on the undifferentiated heavy lifting\u0026rdquo; — is usually read as an argument for maximum abstraction. It is better read as an argument for choosing the abstraction level per workload.\nAWS CDK makes the dial explicit. The construct levels:\nL1 constructs \u0026ldquo;map directly to a single AWS CloudFormation resource\u0026rdquo; and \u0026ldquo;offer no abstraction.\u0026rdquo; L2 constructs \u0026ldquo;include sensible default property configurations, best practice security policies, and generate a lot of the boilerplate code and glue logic for you.\u0026rdquo; L3 constructs, or patterns, are \u0026ldquo;the highest-level of abstraction,\u0026rdquo; bundling resources \u0026ldquo;configured to work together\u0026rdquo; — ApplicationLoadBalancedFargateService is the canonical example. A PaaS is an L3 somebody else wrote and charges for. Most of the time that is exactly what you want. The question is what happens when the L3 stops fitting: you need VPC access to a private database, or a 15-minute function, or IAM-scoped access to a queue, or a region the vendor does not offer.\nIf your only skill is the PaaS, the answer is \u0026ldquo;migrate.\u0026rdquo; If you have CDK and Terraform fluency on the same substrate, the answer is \u0026ldquo;drop one level.\u0026rdquo; Same AWS account, same region, same IAM model, one more construct.\nWhy CDK and Terraform, specifically #Two tools, because they are good at different things and you will encounter both.\nCDK is the right tool when the infrastructure is tightly coupled to application code: a Lambda handler and its event source and its IAM policy in one TypeScript file, type-checked together. Its L2 constructs encode AWS\u0026rsquo;s own opinions about least-privilege and sensible defaults, which is worth a lot when you are moving fast.\nTerraform is the right tool when the infrastructure outlives any single application, spans providers, or is owned by a platform team that needs plan output a reviewer can read. The AWS provider is the most widely used provider in the registry, and its resource model is the one every cloud engineer you hire will already know.\nCDK for Terraform closes the gap for teams that want CDK\u0026rsquo;s programming model with Terraform\u0026rsquo;s state and provider ecosystem. AWS CDK\u0026rsquo;s construct model is portable: the docs note that constructs \u0026ldquo;are available to use with other tools such as CDK for Terraform (CDKtf), CDK for Kubernetes (CDK8s), and Projen.\u0026rdquo;\nFluency in both means you can read any infrastructure repository you are handed and contribute to it the same week. That is the practical definition of flexible.\nThe market makes the bet safe #Synergy Research Group\u0026rsquo;s Q4 2024 numbers: a $330 billion market, with Amazon at 30%, Microsoft at 21%, and Google at 12%. AWS is the largest single substrate, and the multi-cloud vendors — Snowflake, Atlas, Confluent — ship on it first.\nThe primitives also rhyme across providers. Regions, availability zones, virtual networks, identity-scoped permissions, managed queues, object storage: learn them deeply on one provider and the others are a vocabulary exercise. Learn them shallowly through a PaaS and you have learned the PaaS.\nHonest counterpoints #Not everything is AWS. Fly.io runs on hardware it operates itself. Cloudflare Workers run on Cloudflare\u0026rsquo;s network. 37signals left AWS entirely and expects to save \u0026ldquo;at least $1.5 million per year by owning our own hardware.\u0026rdquo; On those platforms AWS fluency transfers less, and DHH\u0026rsquo;s argument — that the hyperscaler premium is itself a cost the abstraction hides — is fair.\nMulti-cloud vendors genuinely abstract the provider. If your team lives inside Snowflake or Kafka semantics, specialization in the vendor\u0026rsquo;s product pays more than specialization in any single cloud underneath it. The underlying region\u0026rsquo;s failure modes are mostly the vendor\u0026rsquo;s problem.\nRenting someone else\u0026rsquo;s operational learning is the point. Vercel built Fluid compute so its customers never had to understand Lambda\u0026rsquo;s concurrency model. Heroku, not its customers, owned the S3-dependency fix in 2017. Most teams should not reproduce that work, and this post is not arguing they should. It is arguing that the engineer who could reproduce it is the one who knows when the rental is worth the price.\nThe thesis #Specialize in AWS. Go deep enough that you can build the L3 yourself, even if you mostly choose not to. Learn CDK and Terraform well enough to drop a level without changing providers. Then use whatever platform fits the workload, knowing what it is made of, what it will cost, and exactly where the exit is.\nThat is not narrowness. That is the only kind of flexibility that survives contact with an incident.\n","date":"11 March 2025","permalink":"https://blog.rickgwaterman.com/posts/its-all-just-aws-under-the-hood/","section":"Posts","summary":"Vercel, Netlify, Supabase, Heroku, Render, PlanetScale: open their region docs and you find AWS region codes. The platforms are thin layers over primitives you can learn directly. Deep AWS specialization — with real fluency in CDK and Terraform — is not narrowing. It is the thing that lets you use any of them, and leave any of them.","title":"It's All Just AWS Under the Hood"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/paas/","section":"Tags","summary":"","title":"Paas"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/adr/","section":"Tags","summary":"","title":"Adr"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/architecture/","section":"Tags","summary":"","title":"Architecture"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/engineering-culture/","section":"Tags","summary":"","title":"Engineering-Culture"},{"content":" Generative AI assisted. The ideas and opinions here are my own; an AI model helped draft and edit the text. The old line is \u0026ldquo;if you\u0026rsquo;re not part of the solution, you\u0026rsquo;re part of the problem.\u0026rdquo; For software and cloud architects the version that matters is narrower: if you are not part of the implementation, you are part of the problem.\nNot the whole implementation. Not the critical path. But a real, committed, reviewed piece of the system you designed, so that the consequences of your decisions land on you before they land on the team.\nThis is not a new idea. It is a twenty-year-old consensus that organizations keep failing to act on.\nThe feedback loop is the job #Gregor Hohpe\u0026rsquo;s Architect Elevator names the failure mode exactly: \u0026ldquo;A well-known architecture department anti-pattern is the \u0026lsquo;ivory tower\u0026rsquo;: architects sit in the penthouse to define how developers should design and build software, without developing any software themselves. Such a setup has one cardinal flaw: it doesn\u0026rsquo;t provide feedback to the architects as to the effectiveness nor the cost of their decisions.\u0026rdquo;\nHe goes further: \u0026ldquo;Worse yet: some architects quite enjoy themselves not having to deal with those consequences.\u0026rdquo;\nThat is the uncomfortable part. An architecture role that never touches the build is comfortable precisely because nothing ever proves it wrong. The diagram was correct. The team must have implemented it badly. A design that cannot be falsified is not a design; it is a position.\nWerner Vogels made the operations version of this argument for Amazon in 2006: \u0026ldquo;The traditional model is that you take your software to the wall that separates development and operations, and throw it over and then forget about it. Not at Amazon. You build it, you run it.\u0026rdquo; His reason was not ideology. \u0026ldquo;Giving developers operational responsibilities has greatly enhanced the quality of the services.\u0026rdquo; The same mechanism applies one level up. You design it, you build part of it.\nAstronauts have been with us since 2001 #Joel Spolsky described the pattern in Don\u0026rsquo;t Let Architecture Astronauts Scare You: \u0026ldquo;When you go too far up, abstraction-wise, you run out of oxygen.\u0026rdquo; And: \u0026ldquo;It\u0026rsquo;s very hard to get them to write code or design programs, because they won\u0026rsquo;t stop thinking about Architecture.\u0026rdquo;\nHis 2008 follow-up gives the tell: \u0026ldquo;The hallmark of an architecture astronaut is that they don\u0026rsquo;t solve an actual problem… they solve something that appears to be the template of a lot of problems.\u0026rdquo;\nThe modern astronaut does not write white papers about XML. They produce a reference architecture in a slide deck, a set of \u0026ldquo;guardrails\u0026rdquo; nobody can find in a repo, and a platform roadmap that assumes an engineering team that does not exist. The slide deck is not wrong. It is unfalsifiable, which is worse.\nTwo kinds of architect #Martin Fowler\u0026rsquo;s Who Needs an Architect? from 2003 splits the role in two.\nArchitectus Reloadus \u0026ldquo;is the person who makes all the important decisions. The architect does this because a single mind is needed to ensure a system\u0026rsquo;s conceptual integrity, and perhaps because the architect doesn\u0026rsquo;t think that the team members are sufficiently skilled to make those decisions.\u0026rdquo;\nArchitectus Oryzus looks different: \u0026ldquo;In the morning, the architect programs with a developer, trying to harvest some common locking code. In the afternoon, the architect participates in a requirements session.\u0026rdquo; Fowler\u0026rsquo;s summary is that \u0026ldquo;the most noticeable part of the work is the intense collaboration,\u0026rdquo; and he offers a better job title: \u0026ldquo;guide, as in mountaineering,\u0026rdquo; someone who \u0026ldquo;is always there for the really tricky stuff.\u0026rdquo;\nHis sharpest line is the one I keep coming back to: \u0026ldquo;an architect\u0026rsquo;s value is inversely proportional to the number of decisions he or she makes.\u0026rdquo; The hands-on architect is not the one who decides everything. It is the one who is close enough to the code to know which decisions actually matter, and takes the hardest one personally.\nThe profession already agreed #Simon Brown, 2010: \u0026ldquo;an architect that codes is more effective and happier than an architect that watches from the sidelines.\u0026rdquo; On the policy some organizations have of keeping architects out of code because they are too valuable: \u0026ldquo;why let your architects put all that effort into defining the architecture if you\u0026rsquo;re not going to let them contribute to its successful delivery?\u0026rdquo;\nBrown again, 2020: \u0026ldquo;Appreciating that you\u0026rsquo;re going to be contributing to the coding activities often provides enough incentive to ensure that your designs are grounded in reality too.\u0026rdquo;\nMark Richards and Neal Ford, in Fundamentals of Software Architecture: \u0026ldquo;The most successful architects we know are those who have broad hands-on technical knowledge coupled with a strong knowledge of a particular domain.\u0026rdquo; Their observation about decay is the consequence of losing that: \u0026ldquo;not enough architects focus their energies on continually analyzing existing architectures. As a result, most architectures experience elements of structural decay.\u0026rdquo;\nThoughtWorks put Coding architects in the Adopt ring of the Technology Radar in August 2010 and kept it there through 2012. Adopt means \u0026ldquo;we feel strongly that the industry should be adopting these items.\u0026rdquo; That was twelve years ago.\nArchitecture is now literally code #The strongest reason to be part of the implementation in 2024 is that the architecture is no longer a document describing the system. It is a Terraform module or a CDK stack in a repository, and it is the system.\nThe VPC topology, the IAM boundaries, the queue retry policies, the Lambda concurrency limits, the Aurora failover configuration: none of that lives in a diagram anymore except as a lagging, usually stale, illustration of what the IaC says. If the architect cannot open a pull request against the infrastructure repository, the architect does not own the architecture. Someone else does, and that someone will eventually stop asking.\nMichael Nygard\u0026rsquo;s Architecture Decision Records belong in the same place: \u0026ldquo;We will keep ADRs in the project repository under doc/arch/adr-NNN.md.\u0026rdquo; Without the rationale next to the code, he notes, a later engineer can only \u0026ldquo;blindly accept the decision\u0026rdquo; or \u0026ldquo;blindly change it.\u0026rdquo; An architect who writes ADRs into the repo, and whose IaC commits reference them, is part of the implementation in the most durable way possible.\nWhat \u0026ldquo;part of the implementation\u0026rdquo; means in practice #It does not mean taking the critical-path feature. Richards and Ford call that the bottleneck trap, and they are right: an architect who owns the one module everyone is waiting on stalls the team and stops doing the rest of the job.\nIt means something like this:\nOwn the infrastructure as code. Write the first version of the CDK stack or Terraform module yourself. Review every change to it. This is the piece of the build that most directly encodes your decisions. Take the hard module, not the big one. The idempotency layer, the outbox pattern, the event schema versioning, the authorizer. Small in lines, large in consequence. Fowler\u0026rsquo;s \u0026ldquo;really tricky stuff.\u0026rdquo; Build the proof of concept before the decision, not after. If the ADR says Step Functions over a custom orchestrator, the ADR should link to a working state machine you deployed. Sit in the review queue. Not as a gate. As the person who reads the most pull requests on the team, because that is where you find out what the architecture actually costs. Carry a pager for what you designed. You build it, you run it. Two or three of those, consistently, is enough. Zero of them is the problem.\nHonest counterpoints #Breadth versus depth. Richards and Ford argue architects should \u0026ldquo;focus on technical breadth rather than technical depth.\u0026rdquo; That is correct, and deep ownership of one module can trade away breadth. The resolution is to choose implementation work that forces breadth: IaC touches every service; the review queue touches every module.\nThe penthouse is also the job. Hohpe\u0026rsquo;s elevator goes both directions. An architect who only lives in the engine room fails to connect strategy to delivery. Brown concedes that being hands-on \u0026ldquo;doesn\u0026rsquo;t necessarily mean that you have to get involved in the day-to-day coding tasks,\u0026rdquo; and some weeks there is no time. The point is engagement over time, not a daily line count.\nScale. Past a certain organization size, one person cannot commit to every repository. The answer is not to stop, but to pick the repository whose decisions are hardest to reverse. Fowler\u0026rsquo;s definition of architecture is \u0026ldquo;things that people perceive as hard to change.\u0026rdquo; Go implement in the place that is hardest to change.\nThe test #Ask any architect one question: what did you merge last month?\nIf the answer is a diagram, the organization has an astronaut. If the answer is a Terraform module, an ADR, a proof of concept, and a stack of reviewed pull requests, it has a guide.\nThe architecture will reflect which one it got.\n","date":"14 May 2024","permalink":"https://blog.rickgwaterman.com/posts/if-you-are-not-part-of-the-implementation/","section":"Posts","summary":"An architect who does not ship code has no feedback loop, and an architecture with no feedback loop is a guess with a diagram. Own a real piece of the build — the hard module, the infrastructure as code, the review queue — or accept that the team will route around you.","title":"If You Are Not Part of the Implementation, You Are Part of the Problem"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/leadership/","section":"Tags","summary":"","title":"Leadership"},{"content":" Generative AI assisted. The ideas and opinions here are my own; an AI model helped draft and edit the text. Most infrastructure decisions are not about capability. They are about who carries the pager.\nWhen a team runs its own queue, its own job scheduler, its own Postgres, or its own orchestration layer, it is not buying flexibility. It is buying an operations job that nobody on the team applied for, and it is paying for that job with the hours that were supposed to go into the product. The managed alternative on AWS — SQS, Step Functions, EventBridge Scheduler, Aurora, DynamoDB, Lambda — is usually less interesting and usually the right call.\nThat is the whole argument. The rest of this post is the evidence.\nInnovation tokens are an infrastructure budget #Dan McKinley\u0026rsquo;s Choose Boring Technology gives the clearest framing. \u0026ldquo;Let\u0026rsquo;s say that we all get a limited number of innovation tokens to spend,\u0026rdquo; he writes. \u0026ldquo;These represent our limited capacity to do something creative, or weird, or hard.\u0026rdquo; Spend them on the problem that makes your company different. Do not spend them on a message broker.\nMcKinley\u0026rsquo;s definition of boring is the useful part. Boring does not mean good. \u0026ldquo;It\u0026rsquo;s boring in the sense that it\u0026rsquo;s well understood. It\u0026rsquo;s bad, but you know why it\u0026rsquo;s bad. You can list all of the main ways it will let you down.\u0026rdquo; That sentence describes SQS perfectly. Everyone knows the visibility timeout gotchas, the at-least-once delivery, the 256 KB message cap. Nobody can list the failure modes of the custom job runner a contractor wrote in 2019, because nobody has read it since.\n\u0026ldquo;Adding the technology is easy,\u0026rdquo; McKinley says. \u0026ldquo;Living with it is hard.\u0026rdquo; Managed services move the living-with-it onto someone whose entire business is living with it.\nAWS says the same thing in its own vocabulary #This is not a contrarian position. It is a literal design principle in the Well-Architected Framework, Operational Excellence pillar: \u0026ldquo;Use managed services: Reduce operational burden by using AWS managed services where possible. Build operational procedures around interactions with those services.\u0026rdquo;\nThe Cost Optimization pillar repeats it from the other direction: \u0026ldquo;Stop spending money on undifferentiated heavy lifting.\u0026rdquo; Werner Vogels has been using that phrase since the mid-2000s; the canonical written version is his 2014 Lambda announcement, where he describes what customers do not want to own: \u0026ldquo;provisioning and scaling servers, keeping software stacks patched and up to date, handling fleet-wide deployments, or dealing with routine monitoring, logging, and web service front ends.\u0026rdquo;\nRead that list again and count how many of those items your self-hosted component drags back onto your team.\nThe hidden tax has a name #Google\u0026rsquo;s SRE book calls it toil: work \u0026ldquo;tied to running a production service that tends to be manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows.\u0026rdquo; Their test is blunt: \u0026ldquo;If your service remains in the same state after you have finished a task, the task was probably toil.\u0026rdquo;\nGoogle caps toil at 50% of an SRE\u0026rsquo;s time. Most small teams running their own infrastructure blow through that number without ever measuring it, because toil hides inside \u0026ldquo;quick\u0026rdquo; tasks: rotating a certificate on the broker, bumping a Postgres minor version, clearing a stuck job, adding disk. None of those tasks leave the system better than they found it. All of them were somebody\u0026rsquo;s evening.\nCharity Majors puts the management obligation plainly in On Call Shouldn\u0026rsquo;t Suck: \u0026ldquo;It is engineering\u0026rsquo;s responsibility to be on call and own their code,\u0026rdquo; and \u0026ldquo;It is management\u0026rsquo;s responsibility to make sure that on call does not suck.\u0026rdquo; Her prescription — \u0026ldquo;Closely track how often your team gets alerted. Take ANY out-of-hours-alert seriously\u0026rdquo; — is much easier to honor when the thing paging you is a Lambda function with a dead-letter queue than a hand-rolled scheduler with a cron entry and a prayer.\nBoring includes observable #Well-Architected lists \u0026ldquo;Implement observability for actionable insights\u0026rdquo; and \u0026ldquo;Anticipate failure\u0026rdquo; right alongside \u0026ldquo;Use managed services.\u0026rdquo; That is not a coincidence. Managed services ship with the observability built in:\nSQS publishes queue depth and message age to CloudWatch without configuration. Step Functions keeps a full execution history you can replay. Lambda gives you invocation counts, errors, throttles, and duration by default. DynamoDB exposes consumed capacity and throttling per table. Aurora surfaces Performance Insights without an agent. Werner Vogels\u0026rsquo; Frugal Architect laws, published at re:Invent this year, make the connection explicit in Law IV: \u0026ldquo;Unobserved Systems Lead to Unknown Costs.\u0026rdquo; A custom component is unobserved until somebody instruments it, and somebody rarely does, because instrumenting it was never on the ticket.\nThe Prime Video case is evidence for, not against #Earlier this year Prime Video published a post about moving one monitoring service off Step Functions and Lambda, cutting its infrastructure cost by more than 90%. The post was widely read as \u0026ldquo;serverless failed.\u0026rdquo; It says the opposite if you read the details (InfoQ summary).\nThe service did stream-defect detection on video. The first version used Step Functions to orchestrate Lambda detectors, passing frames through S3. The problem was specific: \u0026ldquo;several state transitions for each second of the analyzed audio/video stream,\u0026rdquo; each billed, plus Tier-1 S3 reads and writes for every intermediate frame. Their fix packed the detectors into a single process on ECS so data could move in memory, and they kept Lambda as the entry point and ECS as the managed runtime.\nSam Newman\u0026rsquo;s read, quoted in DevClass, is that the post \u0026ldquo;is really speaking more about pricing models of functions vs long-running VMs.\u0026rdquo; Adrian Cockcroft\u0026rsquo;s response says the team did what he has advised for years: build serverless first, then \u0026ldquo;optimize serverless applications by also building services using containers to solve for lower startup latency\u0026rdquo; where needed. Jeremy Daly\u0026rsquo;s summary: \u0026ldquo;Serverless First,\u0026rdquo; not \u0026ldquo;Serverless Only.\u0026rdquo;\nThat is the pattern. Start on the managed, boring path. Measure. When one hot path\u0026rsquo;s billing model fights the workload, re-platform that one path onto the next-most-boring managed service. Prime Video never ran their own orchestrator. They moved from one AWS-managed runtime to another.\nWhat custom orchestration becomes #Brian Foote and Joseph Yoder described the end state in 1997, in Big Ball of Mud: a system whose \u0026ldquo;organization, if one can call it that, is dictated more by expediency than design.\u0026rdquo; Their observation about throwaway code is the one that should worry you: it \u0026ldquo;was intended to be used only once and then discarded. However, such code often takes on a life of its own.\u0026rdquo;\nA custom retry loop becomes a custom scheduler becomes a custom state machine with no execution history and one person who understands it. Step Functions imposes that structure up front — states, retries, catch blocks, timeouts — and does so in a way the next engineer can read from the console.\nHonest counterpoints #Pricing models can fight you at scale. Per-transition and per-invocation billing is wrong for chatty, tightly coupled, high-frequency work. Prime Video is the proof. The answer is to measure and move the specific component, not to start on EC2 for everything.\nSteady, predictable load changes the math. DHH\u0026rsquo;s Why we\u0026rsquo;re leaving the cloud argues that at steady scale 37signals was \u0026ldquo;paying an at times almost absurd premium for the possibility that it could\u0026rdquo; burst. He is right about his case, and he concedes the case this post is about: cloud wins \u0026ldquo;when your application is so simple and low traffic that you really do save on complexity by starting\u0026rdquo; there, or when traffic swings. Most teams are in that second category longer than they think.\nLock-in is real. Amazon States Language, DynamoDB\u0026rsquo;s data model, and Lambda event shapes do not port. Gregor Hohpe\u0026rsquo;s Don\u0026rsquo;t get locked up into avoiding lock-in is the right frame: lock-in is a trade-off to price, not a taboo. \u0026ldquo;We may happily accept some amount of lock-in if we get a commensurate pay-off,\u0026rdquo; and \u0026ldquo;additional investment into reducing lock-in actually leads to higher total cost.\u0026rdquo; If the utility is high, accept it and write it down in an ADR.\nThe default #When there is a managed AWS service for the job, use it. When there is not, look harder, because there usually is. When you have confirmed there is not, build the smallest thing that works, instrument it on day one, and treat it as a liability on the balance sheet rather than an asset.\nBoring is not a compromise. It is the thing that lets the interesting work happen.\n","date":"4 December 2023","permalink":"https://blog.rickgwaterman.com/posts/boring-managed-services-win/","section":"Posts","summary":"Every queue, scheduler, or database you run yourself spends an innovation token on something your customers will never notice. Managed AWS services are boring in the best sense: their failure modes are documented, their metrics ship by default, and the toil they remove is the toil nobody measures.","title":"Boring Managed Services Win"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/managed-services/","section":"Tags","summary":"","title":"Managed-Services"},{"content":"","date":null,"permalink":"https://blog.rickgwaterman.com/tags/operations/","section":"Tags","summary":"","title":"Operations"},{"content":"I\u0026rsquo;m Rick Waterman, a Lead Cloud Architect in the Vancouver, WA / Portland, OR metro. I\u0026rsquo;ve spent ten-plus years in backend and cloud engineering — leading architecture and delivery for AWS serverless systems, event-driven services, and company-wide data platforms, and before that building full-stack systems for payment-risk detection at scale, financial research tooling, and agricultural IoT.\nDay to day I care about durable architecture: systems teams can actually operate, CI/CD that gates every merge, observability that catches problems before users do, and infrastructure as code for all of it. I hold the AWS Certified Developer (2023) and AWS Solutions Architect (2020) certifications.\nThis blog is where longer-form writing lands. Working references live at notes.rickgwaterman.com, and the main site with my resume is rickgwaterman.com.\nAI disclosure #The ideas and opinions on this blog are my own. I draft and edit with generative AI assistance, including source research, and review everything before it ships. Each post carries a short notice at the top, and factual claims link to their sources.\n","date":"1 January 0001","permalink":"https://blog.rickgwaterman.com/about/","section":"Blog (Generative AI Assisted)","summary":"","title":"About"}]