Saturday, September 26, 2026

Traditional Locks vs. Lock-Free CAS: Choosing the Right Concurrency Model

TL;DR: Lock-free primitives (CAS) offer near-zero overhead under low contention, but traditional locks (Mutexes) protect your CPU when contention spikes. Here is how to choose between them in high-throughput systems.


The Core Philosophy: Pessimism vs. Optimism

When multiple threads access shared memory, you face a fundamental architectural choice:

  1. Traditional Locks (Mutex / Semaphore): Pessimistic. Assumes conflicts are frequent. Threads lock the critical section before touching memory; contending threads are descheduled to sleep by the OS kernel.
  2. Lock-Free CAS (Compare-And-Swap): Optimistic. Assumes conflicts are rare. Threads compute updates speculatively and apply them via hardware-level atomic instructions (like x86 CMPXCHG), retrying in user-space if another thread modified the value first.

Head-to-Head Comparison

Metric / Dimension Traditional Lock (Mutex / Semaphore) Lock-Free CAS (Atomic Primitives)
Concurrency Model Pessimistic (assumes conflict is likely) Optimistic (assumes conflict is rare)
Thread State on Contention Blocked / Descheduled (OS-level sleep) Spins in user-space retry loop
Low Contention Overhead Moderate (lock acquisition metadata) Negligible (single CPU instruction)
High Contention Overhead Heavy context-switch latency, but bounded CPU usage CPU burning & cache line thrashing (spin-retry loops)
Failure Modes Deadlocks, Priority Inversion ABA problem, Thread Starvation
Typical Use Cases Multi-variable updates, long workflows, I/O operations Metrics counters, sequence IDs, Disruptor ring buffers, non-blocking queues

When Lock-Free CAS Shines (and When It Backfires)

The Win: Sub-Microsecond Throughput

For single-variable state updates (e.g., atomic counters, rate-limit sequence generators, metrics collection) under moderate concurrency, CAS incurs almost zero overhead because it executes in user-space without kernel transitions.

// Fast, non-blocking sequence generation
public long nextId(AtomicLong counter) {
    return counter.incrementAndGet(); // Single atomic CAS loop
}

The Trap: The Contention Cliff

When dozens of threads slam into the same memory address simultaneously, CAS threads spin in a tight retry loop: * Cache Line Bouncing: Invalidation storms flood the CPU cache coherency bus (MESI/MOESI protocol). * CPU Saturation: CPU cores spike to 100% running empty retry cycles, causing tail latencies to explode.


When to Stick with Traditional Locks

Mutexes carry the overhead of kernel transitions, but they offer critical safety guarantees:

  • Compound Invariants: Modifying two related variables atomically (e.g., transferring funds between account A and account B).
  • I/O & Long-Running Critical Sections: If a critical section involves disk access, network calls, or sleep, a lock yields the CPU core to other productive threads instead of burning CPU cycles.
  • Controlled CPU Under Pressure: Under extreme burst traffic, sleeping threads preserve system resources rather than crashing cache subsystems.

Quick Decision Checklist

  • Choose Lock-Free (CAS / Atomics) if:

    • You are updating isolated variables or lightweight non-blocking queues (e.g., LMAX Disruptor).
    • Lock hold duration is measured in nanoseconds.
    • You need to avoid thread parking latency in ultra-low-latency paths.
  • Choose Traditional Locks (Mutex / Semaphore) if:

    • Operations involve multiple variables or multi-step business transactions.
    • Contention is heavily saturated and you need predictable CPU bounds.
    • Work inside the critical section includes blocking operations or I/O.

SSO vs SAML vs OAuth vs OIDC — And the Enterprise Identity Stack Behind Them

A fast-growing company hits 500 engineers. IT is drowning in Jira tickets: "reset my password," "I need access to Datadog," "why can't the new contractor see the staging AWS account?" The VP of Engineering wants one login for everything. The security team wants audit trails. The platform team wants to stop hand-configuring Okta.

Every identity protocol and system in this post exists because that company's pain is real, and each concept solves a specific piece of it. Let's walk through what happens when you actually automate enterprise identity — from the first SSO rollout to full identity-as-code.


Phase 1: "Everyone Is Drowning in Passwords" — SSO

The first ask is always the same: "Can people just log in once?"

SSO (Single Sign-On) is a UX pattern — not a protocol. The user enters credentials once at a central Identity Provider (IdP) and gets access to Slack, Jira, GitHub, Datadog, and AWS without logging in again. That's the outcome. The question is how.


Phase 2: "Wire Up the Enterprise Apps" — SAML

The platform team opens Salesforce, Workday, and Jira's admin consoles. Every one of them has the same integration option: SAML 2.0.

SAML (Security Assertion Markup Language) is an enterprise authentication and federation protocol based on XML. The IdP (Okta, in this case) signs XML "assertions" — digitally signed documents that say "this user is jane@acme.com, she authenticated at 2:14 PM, and she belongs to the Engineering group." The Service Provider (Salesforce, Jira) trusts that assertion and lets her in.

SAML dominates traditional corporate SSO because enterprise SaaS vendors have supported it for 15+ years. If the vendor's admin console has an "SSO" tab, it's almost certainly SAML.

What it answers: "Who is this user?" — authentication and identity federation.


Phase 3: "Let the Scheduling Tool Read My Calendar" — OAuth

An engineer wants to connect a meeting-scheduling app to their Google Calendar. The app doesn't need to know who the engineer is — it just needs permission to read calendar events on their behalf.

OAuth 2.0 (Open Authorization) is an authorization framework — not an authentication protocol. It does not verify who you are; it answers what can this app do on my behalf? OAuth issues access tokens that grant limited, scoped, delegated access to specific resources — without exposing the user's password.

This is the pattern behind every "Allow this app to access your Google Calendar?" consent screen.

What it answers: "What can this app do on my behalf?" — delegated access, not identity.


Phase 4: "Build a Modern Login for Our Customer Portal" — OIDC

Now the product team wants a customer-facing portal with "Log in with Google" and "Log in with Apple." SAML is too heavyweight for mobile and SPAs. OAuth alone doesn't tell you who logged in. They need both authentication and authorization in one modern, lightweight flow.

OIDC (OpenID Connect) is an authentication layer built on top of OAuth 2.0. It adds an ID token — a JSON Web Token (JWT) containing the user's identity (email, name, profile picture) — alongside the OAuth access token. One flow, two tokens: the ID token proves who they are, the access token proves what they can touch.

"Log in with Google" is an OIDC flow. So is "Log in with Apple," "Log in with GitHub," and every modern consumer SSO integration.

What it answers: "Who is this user?" + basic profile + "What can they access?" — authentication and authorization.


How the Protocols Compose

At this point the company is running all three, and they aren't competing — they're layered:

  1. SSO is the goal — one login for all tools.
  2. SAML delivers enterprise app login — Salesforce, Workday, Jira.
  3. OIDC delivers modern app login — the customer portal, mobile apps.
  4. OAuth 2.0 runs underneath OIDC and handles API delegation when an app needs to call APIs on the user's behalf after login.
SAML 2.0 OAuth 2.0 OpenID Connect (OIDC)
Primary function Authentication / Federation Authorization (Access) Authentication + Authorization
Answers "Who is this user?" "What can this app do on my behalf?" "Who is this user?" + basic profile
Data format Verbose XML JSON / Bearer Token JSON Web Tokens (JWT)
Best use case Enterprise corporate apps Delegated API access Modern web, mobile, & consumer login
Typical tools Okta, ADFS, PingFederate Google/GitHub OAuth apps Okta OIDC apps, Auth0

Phase 5: "Where Does Everyone's Account Actually Live?" — IdP and Identity Store

SSO is working. But then: "We acquired a company. Their engineers are in Active Directory. Ours are in Okta. How do we merge?"

This is where the IdP and Identity Store distinction matters:

  • IdP (Identity Provider): The system that authenticates users and issues tokens — Okta, Azure AD, PingFederate. It answers: "Who authenticated this person?"
  • Identity Store: The database of record for who exists — Active Directory, Okta Universal Directory, LDAP. It answers: "Where does identity data actually live?"

Okta can be both the IdP and the identity store. But in many enterprises, Okta authenticates users while Active Directory remains the source of truth for employee records. Understanding which system is authoritative for identity data is critical when you start automating provisioning.


Phase 6: "New Hires Wait 3 Days for Access" — SCIM

The company is hiring 20 engineers a month. HR creates the employee in Workday. IT manually creates accounts in Okta, Jira, GitHub, Datadog, and AWS. It takes three days and someone always misses a system.

SCIM (System for Cross-domain Identity Management) is a protocol for automatically syncing user and group data between systems. When HR creates a user in Workday, SCIM pushes that identity downstream: Okta account created, group memberships assigned, app access provisioned — all within minutes, no tickets.

When someone leaves? SCIM deprovisioning removes their accounts everywhere. No orphaned credentials.

What it answers: "How do accounts get created, updated, and removed downstream?" — automated lifecycle management.


Phase 7: "Not Everyone Should See Everything" — RBAC

With 500 people accessing 40 apps, the security team asks: "Why does every engineer have admin access to the production AWS account?"

RBAC (Role-Based Access Control) ties permissions to roles, not individuals. An engineer gets the platform-eng role, which grants read-only access to production, write access to staging, and admin access to dev. A new hire gets the role; a departing engineer loses it. Permissions flow from the role, not from individual grants scattered across 40 apps.

What it answers: "What is this role allowed to do?" — access governance at scale.


The Full Stack in One View

Every concept is a layer, and they compose:

SAML OAuth 2.0 OIDC SSO IdP Identity Store RBAC SCIM
What it is XML auth/identity protocol Delegated authorization protocol Auth layer on top of OAuth 2.0 UX pattern: one login, many apps System that authenticates users and issues tokens Database of record for identities Access-control model: permissions tied to roles Protocol for syncing user & group data between systems
Answers "Who is this user?" "What can this app do on my behalf?" "Who is this user?" + basic profile "Is the user already logged in?" "Who authenticates users for my apps?" "Where does identity data actually live?" "What is this role allowed to do?" "How do accounts get created/updated/removed downstream?"
Typical tool Okta, ADFS, PingFederate Google/GitHub OAuth apps Okta OIDC apps, Auth0 Okta dashboard, Azure AD portal Okta, Azure AD, Google Workspace Active Directory, Okta Universal Directory, LDAP Okta groups + app permissions, AWS IAM roles Okta SCIM connectors, Provisioning Agents

Phase 8: "Stop Clicking Through Admin Consoles" — Identity as Code

The company now has 2,000 engineers, 120 SaaS apps, 30 AWS accounts, and 4 Okta admins. Every access change is a manual console click. An audit reveals 47 orphaned service accounts, 12 users with permissions they shouldn't have, and zero rollback capability.

The platform team's mandate: treat identity like infrastructure — define it in code, review it in PRs, deploy it through pipelines.

Admin-Console IAM IAM-as-Code
Change process Click through Okta/AWS console terraform apply from a reviewed PR
Audit trail Console logs, screenshots Git history — who approved what, when
Rollback Manual, error-prone git revert + terraform apply
Drift detection Hope and periodic manual review terraform plan shows exact drift
Access reviews Spreadsheet + calendar reminder GitOps: access defined in code, PR-based review cycles
Scale One admin, dozens of apps One pipeline, hundreds of apps and policies

What this looks like in practice: - Okta app assignments and group rules declared in Terraform (via the okta provider) - AWS IAM policies, roles, and permission boundaries managed as .tf files - SCIM provisioning configured as code — new hire onboarding triggers downstream account creation automatically - Access reviews run as PR-based workflows: remove an engineer from a group in code, the pipeline deprovisions everywhere


Phase 9: "Our CI/CD Pipeline Has a Static AWS Key" — Workload Identity

The security team finds a long-lived AWS access key hardcoded in a GitHub Actions workflow. It has AdministratorAccess. It was created two years ago. Nobody knows who owns it.

Workload identity is identity for non-human principals — services, CI/CD pipelines, and AI agents. The solution: OIDC-based workload identity federation. GitHub Actions presents an OIDC token to AWS, which exchanges it for short-lived STS credentials scoped to exactly the permissions that pipeline needs. No static secrets. No keys to rotate. No keys to leak.

Concept What It Is Why It Matters
AWS IAM AWS's native identity/permission system — users, roles, policies The thing you push policy changes into via Terraform
AWS IAM Identity Center (formerly AWS SSO) AWS's broker that federates an external IdP (Okta) into AWS account access via SAML/OIDC Classic pattern: Okta (IdP) → Identity Center → temporary AWS role credentials
IAM Roles vs Users Roles are assumed temporarily (via federation/STS); users have persistent credentials Platform-as-code work favors roles — no static creds to rotate or leak
Workload / machine identity Identity for services, CI/CD pipelines, or AI agents — not humans Non-human principals are the frontier problem in IAM right now
Terraform + IAM Declaring IAM policies, Okta apps, and group assignments as .tf code instead of console clicks This is the actual deliverable of "identity as code"

Decision Framework: Picking the Right Combination

Building... Use
Consumer-facing app (web/mobile login) OIDC (+ OAuth 2.0 for API access)
Enterprise B2B tool with corporate SSO SAML 2.0 for authentication, SCIM for provisioning
API-to-API delegation (no user present) OAuth 2.0 client credentials flow
Internal platform managing identity at scale OIDC + SCIM + RBAC + Terraform, all as code
Non-human / workload identity (CI/CD, AI agents) OIDC-based workload identity federation — no static secrets

Key Takeaways

  1. SSO is an outcome, not a protocol — SAML and OIDC are the protocols that deliver it.
  2. OAuth ≠ authentication — it handles authorization (delegated access). OIDC adds the identity layer on top.
  3. SAML dominates enterprise, OIDC dominates consumer/mobile — know when to use which.
  4. SCIM eliminates the 3-day onboarding wait — automated provisioning and deprovisioning across all downstream systems.
  5. RBAC governs access at scale — permissions tied to roles, not individuals, across hundreds of apps.
  6. Identity as code is the endgame — Terraform + GitOps for IAM policies, Okta apps, and access reviews eliminates console drift and scales to thousands of engineers.
  7. Workload identity is the frontier — non-human principals (services, pipelines, AI agents) need short-lived tokens via OIDC federation, not static secrets.

References: ISDecisions, Dev.to — Neelendra Tomar, Nikki Siapno — LinkedIn, Okta, Auth0, Microsoft Learn, OneLogin, Pomerium, Fortinet, Oloid

Tuesday, August 18, 2026

Getting Started with CockroachDB HNSW Vector Search: A 5-Minute Guide

TL;DR: CockroachDB now supports vector search directly within its distributed SQL engine using HNSW (Hierarchical Navigable Small World) indexes. This guide provides a quick "Hello World" SQL setup for cosine distance indexing and explores practical enterprise use cases where distributed vector search excels.


Why Vector Search in CockroachDB?

Until recently, building AI-powered applications meant running a dual-database architecture: a relational database (PostgreSQL, MySQL, CockroachDB) for transactional app data, and a separate vector database (Pinecone, Qdrant, Milvus) for embeddings.

This architecture introduces synchronization lag, dual-write failure risks, complex ETL pipelines, and fragmented governance.

CockroachDB solves this by integrating pgvector-compatible vector search natively into its distributed, multi-region SQL database. You get ACID transactions, global resilience, and horizontal scaling alongside fast Approximate Nearest Neighbor (ANN) vector queries.


Core Concepts in 60 Seconds

  1. VECTOR(d) Data Type: Stores float arrays representing high-dimensional embeddings (e.g., 1,536 dimensions from OpenAI text-embedding-3-small or 768 dimensions from Google Gemini embeddings).
  2. HNSW Index: Hierarchical Navigable Small World is a graph-based indexing algorithm that organizes vectors into multi-layer graphs. It provides sub-linear query time with ultra-fast nearest-neighbor lookups.
  3. Cosine Distance (vector_cosine_ops): Measures the angular distance between vectors regardless of magnitude (normalized range $0$ to $2$). In CockroachDB SQL, cosine distance uses the <=> operator (where $0$ indicates identical direction).

Hello World: Step-by-Step Example

Let's set up a simple product recommendation system based on 3-dimensional embeddings.

Step 1: Create the Table

CREATE TABLE products (
    id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    name STRING NOT NULL,
    category STRING NOT NULL,
    price DECIMAL(10, 2) NOT NULL,
    embedding VECTOR(3) NOT NULL
);

Step 2: Create the HNSW Cosine Index

Build an HNSW index specifically tuned for cosine distance queries:

CREATE INDEX idx_products_embedding_hnsw 
ON products 
USING hnsw (embedding vector_cosine_ops);

Note: CockroachDB automatically constructs the HNSW multi-layer graph for fast similarity retrieval.

Step 3: Insert Sample Embeddings

Insert products with simulated normalized 3D vectors:

INSERT INTO products (name, category, price, embedding) VALUES
    ('Wireless Noise-Canceling Headphones', 'Electronics', 299.99, '[0.91, 0.38, 0.17]'),
    ('Bluetooth Ergonomic Earbuds',       'Electronics', 129.99, '[0.88, 0.42, 0.21]'),
    ('Mechanical Gaming Keyboard',       'Electronics', 159.99, '[0.45, 0.85, 0.26]'),
    ('Ergonomic Mesh Office Chair',      'Furniture',   349.00, '[0.12, 0.31, 0.94]'),
    ('Standing Adjustable Desk',         'Furniture',   499.00, '[0.15, 0.28, 0.95]');

Step 4: Query Top-K Nearest Neighbors

To search for products similar to a query vector [0.90, 0.40, 0.18] (e.g., an audio gear search query):

SELECT 
    name, 
    category, 
    price,
    embedding <=> '[0.90, 0.40, 0.18]' AS cosine_distance
FROM products
ORDER BY embedding <=> '[0.90, 0.40, 0.18]'
LIMIT 3;

Expected Output:

name category price cosine_distance
Wireless Noise-Canceling Headphones Electronics 299.99 0.0003
Bluetooth Ergonomic Earbuds Electronics 129.99 0.0012
Mechanical Gaming Keyboard Electronics 159.99 0.1341

The <=> operator computes the cosine distance. Ordering by cosine distance ascending yields the most semantically relevant items first!


Practical Real-World Applications

1. Unified Hybrid Search (SQL Filtering + Vector Similarity)

In pure vector databases, filtering by structured attributes (e.g., price < 200 AND category = 'Electronics') is inefficient or requires complex pre/post-filtering.

With CockroachDB, you can run single-query hybrid searches backed by ACID transactions:

SELECT id, name, price, embedding <=> '[0.90, 0.40, 0.18]' AS distance
FROM products
WHERE category = 'Electronics' AND price <= 200.00
ORDER BY embedding <=> '[0.90, 0.40, 0.18]'
LIMIT 5;

2. Multi-Tenant Enterprise RAG (Retrieval-Augmented Generation)

For AI SaaS products serving thousands of enterprise clients, data isolation is crucial. CockroachDB allows you to store document chunks and embeddings alongside tenant metadata in partitioned tables:

SELECT chunk_text, document_id
FROM knowledge_base_chunks
WHERE tenant_id = 'tenant_12345'
ORDER BY chunk_embedding <=> $query_embedding
LIMIT 5;

This guarantees strict tenant data boundary isolation while leveraging CockroachDB's distributed horizontal scalability.

3. Real-Time Fraud & Anomaly Detection

Financial institutions generate embeddings for user transactions or access patterns. By comparing incoming events against known fraud vector clusters using HNSW cosine distance: - Match distance $< 0.05$ flags instant high confidence anomalies. - Transaction processing occurs in the same distributed database handling account balances, preventing split-brain inconsistencies.

Search product catalogs, audio clips, image embeddings (CLIP), or code repos without exact keyword matches. Cosine similarity focuses on directional alignment of embedding spaces, making it ideal for normalized semantic vectors produced by modern LLMs and vision models.


Key Takeaways

  • HNSW + Cosine Distance (<=>) delivers high-throughput, low-latency similarity search for normalized AI embeddings.
  • No separate vector DB needed: Consolidate transactional state (Postgres/CockroachDB SQL) and vector embeddings into one unified storage engine.
  • Distributed & Resilient: Scales across multiple regions and nodes seamlessly, with zero-downtime schema changes and multi-master fault tolerance.

Get started by trying vector queries in CockroachDB v24.2+ or CockroachDB Cloud!

Friday, August 14, 2026

How Roblox's Cache Sustained 1.38B QPS Beyond Redis Limits

Roblox's caching layer handles 1.38 billion QPS across 6,000+ Redis nodes. When Grow a Garden drove a 10x traffic surge, they survived on existing hardware with a federated architecture and smart efficiency wins.


The Redis Ceiling and How They Broke It

Redis clusters hit a practical limit at 400–500 nodes — beyond that, the Gossip protocol's internode chatter consumes too much CPU and network. AWS ElastiCache enforces this same cap.

Roblox's largest backend service needed 100 million QPS from a single logical cluster (up from 10M in three months). Their solution: a federated architecture — a client reverse proxy in front of multiple independent 400-node Redis clusters, presented as one unified service. Today that's 6,000+ nodes across 15+ clusters, scaling horizontally by simply adding more.


Zero-Downtime Migration

Migrating live production clusters used a three-stage pipeline — no application changes required:

  1. Dual-write to both old and new clusters simultaneously
  2. Read switch once TTL-bounded data reaches parity
  3. Decommission the original cluster

The proxy layer absorbed all migration complexity transparently.


Surviving a 10x Surge on Bare Metal

Grow a Garden shattered gaming CCU records — from 2.8M to 21M concurrent users in three months. Impact: 10x traffic to the largest Redis cluster, all on on-prem hardware with no cloud burst option.

They optimized both their memory-bound Redis fleet (better scheduling, reduced fragmentation) and compute-bound Envoy proxies (tighter autoscaling, leaner health checks). The biggest win: colocating both workloads on the same machines with cgroup isolation — memory-heavy Redis alongside CPU-heavy Envoy — cutting capacity needs by 25%.


Reliability at Scale

  • Hot key detection: Sampling live traffic to throttle overloaded keys before they cascade
  • Partial failure tolerance: Batched requests return partial results instead of all-or-nothing failures
  • Chaos engineering: Regular fault injection (node/rack outages, latency) in production
  • Client-side guardrails: Rate limiting and retry budgets prevent thundering herds

Key Takeaways

  1. Federate when a single component can't scale — proxy in front of multiple clusters works for caches, databases, and queues alike
  2. Dual-write migrations work beautifully for TTL-bounded data — no complex replication needed
  3. Bin-pack complementary workloads — memory-bound + CPU-bound on the same machines with cgroup isolation is free capacity
  4. Design batch operations for partial failure from day one

Next up: Roblox is building a multitenant caching service on ValKey for better cost efficiency and resource isolation.


Commentary on Roblox's engineering blog post by Sen Li, Anders Persson, Pranish Pantha, and Jeffrey Zhong (March 2026).

Tuesday, May 19, 2026

Why Apache Arrow Eliminates Deserialization: The Power of Columnar Memory-Mapped Data

Every data engineer has felt the pain of deserialization — that invisible tax you pay before your code can do anything useful. But what if the data on disk was already in the format your program needs? That's the radical promise of Apache Arrow IPC.

The Traditional Way: Row-by-Row Deserialization

In row-oriented formats (JSON, CSV, Protobuf, Avro), data is stored like this:

Row 1: {name: "Alice", age: 30, score: 95.2}
Row 2: {name: "Bob",   age: 25, score: 88.7}
Row 3: {name: "Carol", age: 28, score: 91.0}

To use this data, your program must:

  1. Parse each row — find field boundaries, decode types
  2. Allocate a new object/struct per row — heap pressure grows linearly
  3. Copy values into those objects — bytes move from kernel buffers into your application's memory

This is deserialization — converting bytes on disk or wire into usable in-memory structures. It's O(n) work proportional to the number of rows, and it happens before you can do anything useful with the data.

For a 100-million-row dataset, that's 100 million parse-allocate-copy cycles just to get started.

The Columnar Way: Arrow IPC

Arrow flips the model entirely. Instead of storing data row-by-row, it stores data column-by-column as flat, typed arrays in memory:

name_buffer:  ["Alice", "Bob", "Carol"]   ← contiguous bytes
age_buffer:   [30, 25, 28]               ← contiguous int32 array
score_buffer: [95.2, 88.7, 91.0]         ← contiguous float64 array

The critical insight: the on-disk/on-wire format IS the in-memory format. There's no transformation needed.

This isn't just a storage optimization — it's an architectural decision that eliminates an entire class of work.

What "Memory-Mapped" Really Means

When you memory-map an Arrow file, the difference is stark:

Traditional:  disk bytes → parse → allocate → copy → usable data
Arrow:        disk bytes → usable data (same thing)

The OS maps the file directly into your process's virtual address space. The age_buffer on disk is already a valid int32[] array — your code can index into it (age_buffer[2] → 28) with zero parsing. The CPU just does a pointer dereference.

No allocation. No copying. No parsing. The data is simply there.

Why This Matters at Scale

  • Zero-copy reads — No CPU cycles wasted on deserialization
  • Instant startup — A terabyte file is "loaded" in microseconds (the OS handles paging on demand)
  • Cache-friendly — Columnar layout means sequential memory access patterns that modern CPUs love
  • Cross-language — The same Arrow buffers work in Python, Rust, Java, C++, and Go without conversion

The Bottom Line

Traditional formats force you to pay a deserialization toll proportional to your data size. Arrow IPC eliminates that toll entirely by making the wire format and the compute format identical. When your data is already in the shape your CPU needs, the fastest deserialization is no deserialization at all.

Tuesday, May 12, 2026

Cactus to Clouds: My First Big Hike

I am not a hiker. Four weeks before May 2, 2026, I had never done a hike outside my hometown. My group trained me at Mission Peak and North Peak in the Bay Area — four weeks of weekend climbs to get my legs ready. Then we drove down to Palm Springs, rented an Airbnb, and on the night of May 1st I set an alarm for midnight.

Cactus to Clouds is rated one of the hardest day hikes in the world. 10,800 feet of climbing. 20 miles. Starting in the Sonoran Desert and ending on the summit of Mt. San Jacinto at 10,834 feet. I didn't fully understand what I had signed up for until I was already on the trail.


The Start: 12:30am, Palm Springs Art Museum

Eight of us left the Palm Springs Art Museum parking lot at 12:30am. It was dark and warm — desert warm, the kind that doesn't feel dangerous yet. I had my headlamp on, trekking poles in hand, and 6 liters of water: 2 in a hydration bladder on my back and 4 in bottles.

The Skyline Trail starts steep immediately. There are no official trail markers for the first several miles — just white painted dots on rocks that local hikers maintain. In the dark, you hike from dot to dot, scanning with your headlamp. Navigation on this section is genuinely hard. I stayed close to my group and trusted them completely. Without them I would have gone off trail within the first mile.

The first gut check is a set of picnic benches about a mile in. We stopped briefly, checked in with each other, and kept moving. The climbing doesn't relent.


Rescue 1 to Grubbs Notch: The Real Work

By the time we passed Rescue 1 at around mile 2.5, the sun was coming up. The desert opened up below us and the views started. But the trail kept going up — relentlessly, brutally up. Almost 1,000 feet of gain per mile on the Skyline.

Tombstone Rock. Flat Rock. The Traverse. Each landmark felt like a small victory and a reminder of how much was left. The Traverse section — steep, loose, shaded — was where I started to feel the weight of the day. My legs were working. My mind was working harder.

Then Grubbs Notch.

When we pulled ourselves over that notch, other hikers around us said the same thing: "The hard part is over. It's easy from here." My group echoed it. I believed it. I made a decision I would regret.


The Mistake: Long Valley Ranger Station

At the Long Valley Ranger Station — mile 10, elevation 8,600 feet — I dumped my four water bottles. All of them. I kept only the 1 liter remaining in my hydration bladder. My hiking partner did the same. We had just climbed 8,000 feet in 10 miles. We were tired. The bottles felt heavy. And everyone said it was easy from here.

What nobody told us — what I didn't understand from reading any guide — is that "easy" is relative. Easy compared to the Skyline. Not easy in any absolute sense. Long Valley to the summit is another 4.5 miles and 2,200 feet of climbing at altitude, through pine forest, past Round Valley, up to Wellman Divide, and then a rocky scramble to the top. For a first-time hiker at hour 10 of a 20-mile day, it is not easy.

It took me 8 hours from Long Valley to the summit and back to the tram. Eight hours on 1 liter of water.


Running Dry

Somewhere between Long Valley and the summit, my bladder ran empty. I squeezed the hose and got nothing. I kept hiking.

The headache came on gradually — a dull pressure behind my eyes that I tried to ignore. I had pain relievers in my pack and took them. I focused on what had gotten me through the Skyline: baby steps. One foot, then the other. Don't look at the distance remaining. Just move.

Nobody in my group had extra water to share. We were all managing our own reserves. I pushed through dry.


The Summit: 5pm

We reached the summit of Mt. San Jacinto at 5pm — nearly 17 hours after leaving the trailhead. Most hikers had already descended to catch the last tram. The summit was quiet. Just me and my hiking partner and a view that stretched across the entire Coachella Valley and beyond.

I don't have the right words for what I felt up there. Positivity is the closest I can get — a kind of amplified, clear-headed positivity that the altitude and the exhaustion and the views all combined to produce. The breeze was cold and clean. We took pictures.

I thought about my kids and my spouse. They had managed everything at home while I trained and traveled. They had cheered me on and asked for updates. Standing on that summit, I dedicated the hike to them. They made it possible.


The Descent and the Tram

Coming down from the summit to the tram station is 5.5 miles of trail you've never hiked before, on legs that have already done 14.5 miles and 10,800 feet of climbing. It is not a victory lap. It is work.

We reached Mountain Station at 8:30pm. The tram was still running. We bought drinks at the gift shop, sat down, and didn't say much for a while. The tram ride down felt surreal — watching the desert floor rise up to meet us, the same terrain we had climbed in the dark now lit by the last light of the day.

Total time: 20 hours. Start to finish.


What I Learned

Don't dump your water at Long Valley. The right approach — especially for someone like me on their first big hike — is 3 liters from the start to the Long Valley Ranger Station, then refill 3 liters there for the second half. That's it. I carried 6 liters to the ranger station and threw 4 of them away. My partner did the same. Don't do that. Other hikers telling you it's easy from Grubbs Notch are not lying — they're just measuring against the Skyline. Carry the water anyway.

"Easy from here" is not a water strategy. I ran out between Long Valley and the summit. I finished with a headache and no water. That is a serious situation on a remote alpine trail. It could have been much worse.

Baby steps are a real strategy. When the trail felt impossible — at Grubbs Notch, on the Long Valley plateau, on the final rocky scramble — I stopped thinking about the summit and focused on the next step. Just the next step. It works.

Your group is everything. I would not have navigated the Skyline in the dark alone. I would not have had the confidence to attempt this hike without four weeks of training with people who believed I could do it. If you're a first-time big hiker, do this with people who know what they're doing.

Train seriously. Mission Peak and North Peak gave me the leg strength to survive this hike. They did not fully prepare me for 20 miles and 10,800 feet. but gave my group the confidence that I could do C2C.


The Numbers

Date May 2, 2026
Start time 12:30am
Summit time 5:00pm
Tram station 8:30pm
Total time ~20 hours
Distance 20 miles
Elevation gain 10,800 ft
Group size 8 hikers
Water carried 6L (dumped 4L at Long Valley — don't do this)

Would I Do It Again?

Yes. But I'd carry the water.

C2C is genuinely one of the hardest things I've done. It is also one of the most rewarding. The Skyline at dawn, the silence on the summit, the tram ride down — those moments don't fade. If you're considering it, train hard, respect the water, and go with people you trust.


Route reference: hikingguy.com/hikes/cactus-to-clouds — the most thorough C2C guide I found, and the one I used to plan this hike.

Thursday, April 30, 2026

Amazon Quick Desktop: A Practical Guide to Getting Started

The views expressed in this post are my own and do not represent the opinions, positions, or endorsements of my current or any former employer.


Amazon Quick Desktop is a native AI assistant for macOS and Windows, launched in April 2026. It connects to your work tools — Google Workspace, Microsoft 365, Slack, Salesforce, Zoom, Jira, and more — and runs persistently in the background, learning the context of your work over time. This guide covers what it does, how to set it up, and how different teams are using it.


What Is Amazon Quick?

Amazon Quick is an AI assistant designed to work across all the tools you already use. Rather than operating within a single app, it connects your local files, calendar, email, and cloud applications into one place and builds a personal knowledge graph — a living model of your role, priorities, relationships, and projects that gets more useful the longer you use it.

The core idea is to reduce the time spent hunting for information across disconnected systems and replace it with a single interface that can answer questions, take actions, and automate workflows on your behalf.


Getting Started

No AWS account or credit card is required to try it.

  1. Sign up at aws.amazon.com/quick using your email, Google, Apple, Amazon, or GitHub account
  2. Download the desktop app for macOS or Windows
  3. Connect your data sources through the onboarding wizard — Google Workspace, Microsoft 365, Slack, Salesforce, Zoom, and others
  4. Quick begins indexing your files and emails in the background and starts building context

Every new signup includes a free 30-day trial of the Plus plan, which includes full desktop access and expanded agent hours. A free tier is available after the trial.


Core Capabilities

Personal Knowledge Graph

Quick builds a private model of your work — the people you collaborate with, the projects you're involved in, the documents and data you use regularly. This context shapes every answer and action. It's stored locally on your device and is not used to train Amazon's models.

Proactive Intelligence

Rather than waiting for you to ask a question, Quick monitors your connected apps in the background and surfaces what needs attention — an unanswered priority email, a Salesforce deal that hasn't been updated, a document waiting for your feedback. You can ask "What am I missing today?" or "What should I prioritize?" and get answers grounded in your actual work context.

Deep Research

Quick Research is a built-in agent that investigates questions by pulling from your internal documents, the public internet, and third-party datasets simultaneously. It creates a research plan, gathers evidence, and produces a fully cited, exportable report. Reports include clickable citations, version history, and export options for PDF, Word, and custom summary formats.

Document and Dashboard Creation

You can generate deliverables directly within the chat — presentations, spreadsheets, Word documents, PDFs, images, and live dashboards. Dashboards connect to your data sources and update automatically. No need to switch to a separate tool.

Workflow Automation

There are two levels of automation:

  • Quick Flows — natural language-based workflows for repeatable tasks like generating weekly reports, routing approvals, or sending automated briefings. No coding required.
  • Quick Automate (Professional/Enterprise) — complex, multi-step automations across systems, such as syncing data between Salesforce and a data warehouse or orchestrating cross-team processes.

Team Spaces and Custom Agents

Spaces are shared knowledge environments where teams pool documents, data, and AI agents around a project. You can also build custom chat agents configured with specific knowledge sources, personas, and guardrails — for example, an HR policy assistant, a sales pipeline tracker, or a project health monitor — and share them across your team.

Extensions

Quick extends beyond the desktop app into browsers (Chrome, Edge, Firefox) and directly into Microsoft Office apps (Word, Excel, PowerPoint, Teams, Outlook) and Slack, so you can access its capabilities without switching windows.


How Teams Are Using It

HR — Onboarding Automation

At the Austin Amazon Quick User Group meetup in January 2026, attendees built a working HR onboarding workflow in under an hour. The setup: upload an employee handbook, leave policy, performance review guidelines, and onboarding checklist into a Space, then create a Quick Flow that accepts employee questions as input, searches the HR documentation, and returns sourced answers automatically. The flow can be shared with the whole HR team.

Operations — Mystery Shopping Review (Ironside Group / HS Brands)

Ironside Group built a solution combining Amazon Quick and Amazon Bedrock to automate survey review for HS Brands Global, a mystery shop provider. The automated system detected inconsistencies, errors, and potential fraud across large volumes of unstructured survey data. Results: review time per batch dropped from days to seconds, annual review capacity scaled from roughly 50,000 shops to millions, and costs were reduced by approximately 85%.

Insurance — Nightly Reconciliation and Compliance (New York Life)

New York Life's Institutional Life division used Quick to replace a manual reporting process that required pulling multiple reports and waiting on analysts. A single conversational agent now handles structured operational data and unstructured documentation together. Their compliance dashboards moved from static reporting to live, self-service analytics, and nightly reconciliation workflows that previously required manual intervention are now automated with Quick Flows.

Manufacturing / Sales — Pipeline Insights (3M)

3M's sales teams used Quick Flows to automate administrative tasks like generating report summaries and updating records, and Quick's agentic capabilities to synthesize information across sales effectiveness, risks, and pricing from multiple platforms.

Pharma Research — Clinical Trial Site Database (Kitsa)

Kitsa built KScout on Amazon Quick — a database of over 300,000 clinical trial sites across 160+ countries — with a team of fewer than five people. Quick powers autonomous site research, medical literature review, and intelligence report generation.


Privacy and Security

  • Data stays on your device — conversation history, memory, knowledge graph, and file indexes are stored locally and not uploaded to the cloud
  • No model training on your data — AWS does not use your data to train models on any plan, free or paid
  • Write operations require approval — Quick will not send an email, update a record, or take an action without your explicit confirmation
  • Certifications: HIPAA eligible, FedRAMP authorized, SOC 2 audited
  • Open standards: supports Model Context Protocol (MCP), allowing integration with third-party agents and tools

Plans

Plan Includes Desktop app
Free Chat, Spaces, custom agents, knowledge bases, extensions No
Plus Full desktop app, expanded agent hours Yes
Professional Quick Sight (BI), Quick Automate, AWS data connectivity Yes
Enterprise SSO, advanced governance, region selection, full AWS infrastructure Yes

Every signup includes a free 30-day Plus trial with up to 10 team members. No AWS account or credit card required.


Tips for Getting Started

Based on community feedback from the Amazon Quick User Group:

  • Start with one specific use case rather than trying to connect everything at once. Pick the workflow that costs your team the most time and build that first.
  • Use Spaces to organize knowledge — upload the documents your team references most often and build agents on top of them.
  • Quick Flows are more accessible than they look — the natural language builder means non-technical team members can create and maintain automations without developer support.
  • The desktop app gets more useful over time — the knowledge graph improves as Quick learns your patterns, so the value compounds the longer you use it.

Get started: aws.amazon.com/quick/desktop


Sources: Amazon Quick features · Amazon Quick FAQs · Amazon Quick customers · Austin Quick User Group meetup · About Amazon: Quick desktop launch. Content was rephrased for compliance with licensing restrictions.