Tagged articles

Performance Optimization

2086 articles · Page 1 of 21
21CTO
21CTO
Oct 1, 2026 · Frontend Development

GitHub's 3-Year Migration to CSS Modules Cuts SSR Time 55%

GitHub completed a three-year migration from CSS-in-JS to CSS Modules, reducing server-side rendering time by 55% and component initialization by 25%, using feature flags, visual regression testing, and gradual rollout to migrate 7,760 sx prop usages across 8 engineers and later GitHub Copilot.

CSS ModulesCSS-in-JSFeature Flags
0 likes · 6 min read
GitHub's 3-Year Migration to CSS Modules Cuts SSR Time 55%
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Sep 29, 2026 · Backend Development

SGLang Multi-Hardware Plugin Architecture and Kunlun Chip Adaptation: A 2-Hour Model Upgrade Case Study

This article details SGLang's device plugin mechanism (Platform and Hook) enabling hardware-agnostic inference, the SGLang-Kunlun plugin's cuda-like compatibility layer, a five-step standardized adaptation process with measurable acceptance criteria, and AI-driven Harness automation that reduced DeepSeek V4 Flash version upgrades to two hours while achieving 1.55x operator speedups via dual-stream vector/matrix compute overlap.

AI-assisted DevelopmentCUDA GraphDevice Plugin
0 likes · 22 min read
SGLang Multi-Hardware Plugin Architecture and Kunlun Chip Adaptation: A 2-Hour Model Upgrade Case Study
Xiaolin Talks Programming
Xiaolin Talks Programming
Sep 27, 2026 · Backend Development

Type-Safe SQL with jOOQ in Spring Boot: Code Generation, DSL & Multi-DB Adaptation

This article details practical integration of jOOQ 3.19 with Spring Boot 3.x for type-safe SQL, covering code generation, DSL queries, multi-database adaptation, performance tuning, security practices, and migration lessons learned from real-world complex reporting and multi-database delivery scenarios.

DSLPerformance OptimizationSpring Boot
0 likes · 34 min read
Type-Safe SQL with jOOQ in Spring Boot: Code Generation, DSL & Multi-DB Adaptation
TonyBai
TonyBai
Sep 24, 2026 · Backend Development

One Engineer, 3 Months, 830K Lines: GitHub Rewrites Copilot Runtime in Rust with Copilot

GitHub engineer Stephen Toub details how a single engineer used Copilot to rewrite the Copilot Agent Runtime from TypeScript to Rust in 14 weeks, producing 832K lines of Rust code with 18x latency improvement and 91% memory reduction at a cost of $120k in tokens, while maintaining quality through in-place migration and rigorous testing.

AI-assisted DevelopmentGitHub CopilotPerformance Optimization
0 likes · 20 min read
One Engineer, 3 Months, 830K Lines: GitHub Rewrites Copilot Runtime in Rust with Copilot
JD Tech
JD Tech
Sep 22, 2026 · Backend Development

18 Advanced Java Coding Practices for High-Performance, Low-CPU Systems (Part 2)

This article details 18 advanced Java coding techniques (19-36) for high-performance, low-CPU systems, covering void/null avoidance, nested code flattening, final/static method optimization, primitive type usage, collection pre-sizing, memory management, immutable collections, loop optimization, wrapper caching, exception reuse, stack trace reduction, try-catch narrowing, bitwise operations, regex precompilation, memory alignment, constant-first comparisons, and batch processing.

Coding Best PracticesHigh PerformanceJVM
0 likes · 52 min read
18 Advanced Java Coding Practices for High-Performance, Low-CPU Systems (Part 2)
php Courses
php Courses
Sep 22, 2026 · Backend Development

PHP Performance Optimization: 6x Speedup from 8.2s to 1.4s with 7 Fixes

A legacy PHP 7.2 project with 8.2s average response time was optimized to 1.4s in two weeks through seven fundamental fixes: enabling OPcache, eliminating N+1 queries via eager loading, adding composite database indexes, implementing Redis caching layers, setting HTTP timeouts for external calls, right-sizing PHP-FPM workers, and separating API middleware from web routes.

Database IndexingLaravelN+1 Query
0 likes · 12 min read
PHP Performance Optimization: 6x Speedup from 8.2s to 1.4s with 7 Fixes
Golang Shines
Golang Shines
Sep 19, 2026 · Operations

SEAL Methodology for Production Troubleshooting: Veteran Ops Toolbox & Case Studies

A 10-year operations veteran shares the SEAL troubleshooting framework (Symptom, Environment, Analysis, Location), a curated toolbox (Prometheus, ELK, perf, tcpdump), real-world case studies (Redis avalanche, MySQL slow queries), incident grading, automation scripts, performance tuning, container/Kubernetes diagnostics, monitoring models, chaos engineering, and AIOps trends.

AIOpsAutomationPerformance Optimization
0 likes · 20 min read
SEAL Methodology for Production Troubleshooting: Veteran Ops Toolbox & Case Studies
JD Tech
JD Tech
Sep 16, 2026 · Backend Development

18 Battle-Tested Java Coding Techniques for High-Performance, Low-CPU Systems (Part 1)

This article presents 18 practical coding techniques derived from million-QPS production systems to reduce CPU consumption and improve performance, covering type unification, string handling, loop optimization, caching strategies, and data structure selection with concrete before/after code examples.

CPU OptimizationCoding Best PracticesGC Reduction
0 likes · 67 min read
18 Battle-Tested Java Coding Techniques for High-Performance, Low-CPU Systems (Part 1)
Spring Full-Stack Practical Cases
Spring Full-Stack Practical Cases
Sep 15, 2026 · Backend Development

Five MapStruct Mistakes That Make Code Harder to Maintain

This article identifies five common MapStruct misuses — complex logic in @Mapping expressions, missing bidirectional relationship handling, over-mapping in hot paths, injecting business services into mappers, and naive PATCH updates — and shows correct Java alternatives for each case.

Code MaintainabilityJPAJava
0 likes · 11 min read
Five MapStruct Mistakes That Make Code Harder to Maintain
SpringMeng
SpringMeng
Sep 14, 2026 · Backend Development

Maven-mvnd: Drop-In Replacement for Faster Java Builds via Daemon & GraalVM

This article introduces Maven-mvnd (mvnd), a drop-in replacement for Apache Maven that uses a persistent daemon process and GraalVM native executables to dramatically accelerate build times, especially for multi-module projects, while maintaining full compatibility with existing Maven POM files and commands.

CI/CDGraalVMJava
0 likes · 9 min read
Maven-mvnd: Drop-In Replacement for Faster Java Builds via Daemon & GraalVM
IT Services Circle
IT Services Circle
Sep 13, 2026 · Backend Development

Hidden Costs of C++ Containers: Why Memory Layout Beats Big-O

This article reveals how C++ STL containers like vector, list, map, and unordered_map incur hidden performance costs through memory layout, cache locality, pointer chasing, and rehash overhead, demonstrating that theoretical time complexity often misleads because cache misses dominate real-world performance.

C++Performance OptimizationSTL
0 likes · 13 min read
Hidden Costs of C++ Containers: Why Memory Layout Beats Big-O
dbaplus Community
dbaplus Community
Sep 13, 2026 · Backend Development

Solving Sharding Routing Latency by Embedding Route Keys in Order IDs

A large OTA platform eliminated sharding query latency by encoding a 4-bit routeKey into 64-bit Snowflake order IDs, enabling direct shard lookup without index table queries, reducing P99 latency from over 1 second to tens of milliseconds and cutting database load by 80% through a three-layer fallback strategy.

Performance OptimizationSnowflakebit-manipulation
0 likes · 20 min read
Solving Sharding Routing Latency by Embedding Route Keys in Order IDs
Xiaolin Talks Programming
Xiaolin Talks Programming
Sep 12, 2026 · Backend Development

Spring Boot SM2/SM3/SM4 Integration: Pitfalls, Key Management & Performance Tuning

This article details practical integration of Chinese national cryptographic algorithms SM2, SM3, and SM4 into Spring Boot, covering dependency setup, encryption/decryption, signing/verification, anti-replay protection, JWT implementation, TLS offloading via Nginx, key rotation strategies, and performance optimizations that raised throughput from 300 to 1300 TPS.

Chinese CryptographyKey RotationMLPS Compliance
0 likes · 20 min read
Spring Boot SM2/SM3/SM4 Integration: Pitfalls, Key Management & Performance Tuning
Woodpecker Software Testing
Woodpecker Software Testing
Sep 11, 2026 · Artificial Intelligence

AI Testing Performance Optimization: Compute, I/O, Scheduling & Observability Deep Dive

This article analyzes performance optimization for AI-driven testing tools across four dimensions—computation, I/O, scheduling, and observability—detailing practical architectural strategies like lightweight models, zero-copy data transfer, dynamic Kubernetes-based scheduling, and multi-layer observability, with real-world case studies showing significant latency and cost reductions.

AI testingKubernetes schedulingPerformance Optimization
0 likes · 9 min read
AI Testing Performance Optimization: Compute, I/O, Scheduling & Observability Deep Dive
CodeOnCode
CodeOnCode
Sep 9, 2026 · Backend Development

Beyond Pause Times: Deep Dive into G1 GC Log Analysis for Java Performance Tuning

This article teaches how to analyze G1 GC logs beyond simple pause times by breaking down phase timings, heap changes, GC causes, and concurrent marking overhead, then mapping log signals like high Update RS or Object Copy to concrete code-level issues such as excessive reference writes, humongous allocations, or weak reference abuse.

G1 GCGC LogsJVM Tuning
0 likes · 27 min read
Beyond Pause Times: Deep Dive into G1 GC Log Analysis for Java Performance Tuning
samdeepthink
samdeepthink
Sep 4, 2026 · Backend Development

HashMap Grouping: From Three Lookups to One computeIfAbsent Call

This article explains how Java 8's computeIfAbsent reduces HashMap grouping from three separate lookups (containsKey, put, get) to a single atomic operation, covering lazy evaluation, concurrency safeguards, null handling, nested grouping patterns, and when to prefer computeIfAbsent over Collectors.groupingBy for streaming data.

HashMapJava 8Map methods
0 likes · 7 min read
HashMap Grouping: From Three Lookups to One computeIfAbsent Call
Xiaolin Talks Programming
Xiaolin Talks Programming
Sep 4, 2026 · Backend Development

Spring Boot High-Concurrency Distributed ID Generation: Snowflake Algorithm, Clock Drift Handling & Performance Optimization

This article walks through implementing Twitter's Snowflake algorithm in Spring Boot for distributed ID generation, covering 64-bit structure, clock drift mitigation strategies (wait, error, historical compensation), pre-generated ID pools with double buffering for 120k QPS, MyBatis-Plus integration, and Meituan Leaf design insights.

MyBatis-PlusPerformance Optimizationclock-drift
0 likes · 17 min read
Spring Boot High-Concurrency Distributed ID Generation: Snowflake Algorithm, Clock Drift Handling & Performance Optimization
Xiaolin Talks Programming
Xiaolin Talks Programming
Sep 2, 2026 · Backend Development

Spring Boot Redis Geo: Nearby Stores & Geofencing – Pitfalls, Optimizations & Production Patterns

A hands-on guide to replacing MySQL spatial queries with Redis Geo for nearby-store search and geofencing, covering Spring Boot integration, GeoHash internals, coordinate-system pitfalls, pagination strategies, cluster key design, and real-world performance numbers from a 5,000-store convenience-chain deployment.

Coordinate SystemsGeoHashGeofencing
0 likes · 20 min read
Spring Boot Redis Geo: Nearby Stores & Geofencing – Pitfalls, Optimizations & Production Patterns
dbaplus Community
dbaplus Community
Aug 30, 2026 · Databases

Why OFFSET Pagination Slows You Down: Switch to Keyset and Cut Latency from Seconds to Milliseconds

The article recounts how an API that originally used OFFSET pagination suffered 2‑3 seconds response time, and after rewriting the query to Keyset (Seek) pagination—adding a secondary sort key for stability and exploring cursor pagination, index‑only OFFSET, and materialized views—the same endpoint now responds in under 200 ms, with detailed benchmark comparisons.

Cursor PaginationDatabase IndexingMaterialized Views
0 likes · 7 min read
Why OFFSET Pagination Slows You Down: Switch to Keyset and Cut Latency from Seconds to Milliseconds
51CTO HarmonyOS Developer Community
51CTO HarmonyOS Developer Community
Aug 28, 2026 · Mobile Development

HarmonyOS Cold Start Optimization: 4-Step Pipeline from Icon Tap to Interactive

This article details a four-step HarmonyOS cold start optimization methodology: baseline measurement, deferred non-critical parsing, taskpool concurrency for heavy parsing, and skeleton screens with in-process caching to decouple first-frame and interactive metrics, validated via dual-track measurement on a sample 'Morning Brief' page.

ArkUICachingHarmonyOS
0 likes · 36 min read
HarmonyOS Cold Start Optimization: 4-Step Pipeline from Icon Tap to Interactive
Woodpecker Software Testing
Woodpecker Software Testing
Aug 28, 2026 · Backend Development

Mastering Concurrent User Testing: Deep Dive into Performance Optimization

The article explains that true concurrent user testing must emulate realistic user sessions with think time and asynchronous actions, identifies three typical bottlenecks—application thread blocking, database connection pool exhaustion, and middleware resource contention—and demonstrates proactive performance‑left‑shift practices and emerging AI‑driven adaptive testing techniques.

AIConcurrent TestingHikariCP
0 likes · 7 min read
Mastering Concurrent User Testing: Deep Dive into Performance Optimization
Woodpecker Software Testing
Woodpecker Software Testing
Aug 27, 2026 · Cloud Native

Adversarial Performance Testing: A Hands‑On Guide to Boost System Resilience

In today’s high‑concurrency, microservice‑driven cloud‑native world, traditional load testing often misses real‑world failure modes, so this guide introduces adversarial performance testing—injecting faults, latency, and malicious traffic—to expose hidden bottlenecks and build resilient systems.

Cloud NativePerformance OptimizationResilience
0 likes · 8 min read
Adversarial Performance Testing: A Hands‑On Guide to Boost System Resilience
IT Learning Made Simple
IT Learning Made Simple
Aug 26, 2026 · Fundamentals

Why Does Cache Hit Rate Make Programs Sometimes Fast and Sometimes Slow?

The article explains how cache hit rate directly impacts program speed, describes the three hit‑rate levels, factors such as data locality, working‑set size and associativity, types of cache misses, and provides practical techniques—data alignment, array traversal, packing, prefetching, and monitoring tools—to improve performance.

C++Performance Optimizationcache
0 likes · 11 min read
Why Does Cache Hit Rate Make Programs Sometimes Fast and Sometimes Slow?
Architect's Guide
Architect's Guide
Aug 26, 2026 · Backend Development

10 Powerful Performance‑Optimization Techniques You Should Try

The article surveys ten practical performance‑optimization tactics—from classic indexing, compression, and caching to prefetching, peak‑shaving, batch processing, and advanced methods such as resource squeezing, horizontal scaling, sharding, and lock‑free designs—explaining their trade‑offs, concrete examples, and when to apply each in real‑world systems.

Batch ProcessingCachingCompression
0 likes · 36 min read
10 Powerful Performance‑Optimization Techniques You Should Try
Ubuntu
Ubuntu
Aug 25, 2026 · Fundamentals

Chinese Engineers Boost Linux 7.3; Linus Uses AI to Fix One‑Char Bug

Linux 7.3’s merge window showcases major performance patches from Chinese teams—ByteDance slashing SMP latency to 1.5 ms, ZTE cutting a 705 ms lock to 1.44 ms, and Xiaomi improving zsmalloc speed up to 1.83×—while Linus Torvalds, aided by Gemini AI, fixed a one‑character bug in the Intel Xe driver, and the release adds features like sched_ext, AMD UALink, Rust on PowerPC, and filesystem throughput boosts, positioning 7.3 as a strong LTS candidate.

AI debuggingLinux kernelPerformance Optimization
0 likes · 10 min read
Chinese Engineers Boost Linux 7.3; Linus Uses AI to Fix One‑Char Bug
Xiaolin Talks Programming
Xiaolin Talks Programming
Aug 25, 2026 · Backend Development

Spring Boot + Neo4j: Building Social Graphs with Modeling, Queries & Performance Optimization

This article demonstrates how to integrate Spring Boot with Neo4j to build a social relationship graph, covering graph data modeling, Cypher query patterns for friend recommendations and shortest paths, performance optimization techniques including indexing and query profiling, and a complete demo implementation for a social feed.

CypherNeo4jPerformance Optimization
0 likes · 25 min read
Spring Boot + Neo4j: Building Social Graphs with Modeling, Queries & Performance Optimization
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 24, 2026 · Artificial Intelligence

Can LLMs Engineer Their Own Infrastructure? A Deep Dive into Φ‑Bench’s Assessment

This article examines Φ‑Bench, a comprehensive LLM infrastructure benchmark that evaluates how well large language models can perform real‑world infra engineering tasks, revealing current models’ strengths, weaknesses, and the gap to becoming true AI engineers.

AI engineeringError AnalysisInfrastructure Benchmark
0 likes · 12 min read
Can LLMs Engineer Their Own Infrastructure? A Deep Dive into Φ‑Bench’s Assessment
ThinkingAgent
ThinkingAgent
Aug 24, 2026 · Artificial Intelligence

Why Quantization and KV‑Cache Are Key to High‑Performance LLM Inference

The article analyzes why the same LLM can exhibit vastly different cost, speed, and concurrency across inference systems, showing that KV‑cache memory management, continuous batching, PagedAttention, quantization trade‑offs, and speculative decoding together determine real‑world throughput and latency.

KV CacheLLM InferencePerformance Optimization
0 likes · 39 min read
Why Quantization and KV‑Cache Are Key to High‑Performance LLM Inference
liandk
liandk
Aug 22, 2026 · Databases

Why Your SQL Is Slow and How Indexes Can Make Queries Run in Seconds

The article explains that indexes are the key to solving most database slow‑query problems, compares full‑table scans with indexed lookups, describes the two main index types in SQL Server, and provides step‑by‑step T‑SQL commands for creating, testing, and removing indexes while warning against over‑indexing.

Database IndexesPerformance OptimizationQuery Tuning
0 likes · 5 min read
Why Your SQL Is Slow and How Indexes Can Make Queries Run in Seconds
Tencent Cloud Developer
Tencent Cloud Developer
Aug 20, 2026 · Artificial Intelligence

Stability Engineering for Large-Scale Distributed Training: Spike Theory in Autonomous Driving

The article analyzes why performance degrades when scaling single‑machine training to thousands of GPUs, attributing it to the straggler effect, exponential spike probability, and system reliability limits, and presents a three‑layer theoretical framework and concrete engineering practices—including HyperAcc, GPU tracing, NUMA binding, and async DataLoader redesign—to keep per‑node spike rates below 0.5 % and achieve stable, 50 % higher throughput in autonomous‑driving model training.

Distributed TrainingGPU scalingHyperAcc
0 likes · 31 min read
Stability Engineering for Large-Scale Distributed Training: Spike Theory in Autonomous Driving
Woodpecker Software Testing
Woodpecker Software Testing
Aug 17, 2026 · Artificial Intelligence

How AI Testing Tools Can Overcome Performance Bottlenecks

The article analyzes why traditional automated test frameworks struggle with micro‑service scale, identifies three hidden sources of AI testing latency, and presents a four‑layer optimization strategy—from data sampling to edge inference—that dramatically improves speed, determinism, and resource elasticity.

AI testingPerformance Optimizationasynchronous pipeline
0 likes · 8 min read
How AI Testing Tools Can Overcome Performance Bottlenecks
Amap Tech
Amap Tech
Aug 14, 2026 · Artificial Intelligence

How AutoSDK Builds a Self‑Evolving AI Coding Loop for Enterprise Delivery

The article explains why a single successful AI‑generated code run is insufficient for enterprise software, and how AutoSDK uses built‑in observability, Loop Engineering, and a four‑stage "observe‑attribute‑intervene‑validate" loop—supported by concrete metrics, trace and log pillars—to achieve stable, continuously improving AI coding delivery.

AI codingLoop EngineeringPerformance Optimization
0 likes · 17 min read
How AutoSDK Builds a Self‑Evolving AI Coding Loop for Enterprise Delivery
liandk
liandk
Aug 14, 2026 · Databases

Hands‑On MySQL Slow Query: Enable Logs, Analyze SQL, and Optimize Performance

The article explains what MySQL slow queries are, why they must be detected, when to enable slow‑query logging, step‑by‑step commands to configure the log, how to simulate and analyze problematic SQL with EXPLAIN, and practical optimization techniques—including index creation and Spring Boot integration—to eliminate performance bottlenecks.

MySQLPerformance OptimizationSQL
0 likes · 10 min read
Hands‑On MySQL Slow Query: Enable Logs, Analyze SQL, and Optimize Performance
360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
Aug 12, 2026 · Databases

ZestKV Achieves 2.3× SET Write Throughput, Surpassing pika3.5 at 700k QPS

By parallelizing both the write side and the response side—using multi‑queue write concurrency, overlapping network and CPU work, batch wake‑ups, and connection‑based sharding—ZestKV raises SET request throughput from 310 k to 700 k QPS, more than double pika 3.5, while preserving full consistency and durability guarantees.

ConsistencyParallelismPerformance Optimization
0 likes · 12 min read
ZestKV Achieves 2.3× SET Write Throughput, Surpassing pika3.5 at 700k QPS
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Aug 6, 2026 · Cloud Native

How ACK Pro Provisioned Control Plane Eliminates Kubernetes Control‑Plane Bottlenecks for Large‑Scale Clusters

ACK Pro introduces a provisioned control‑plane mode that replaces reactive scaling with preset performance tiers, guaranteeing deterministic capacity for thousands of nodes and tens of thousands of Pods, and a real‑world AI training case shows reduced pod‑startup latency, eliminated HTTP 429 errors, and about 30% faster training cycles.

ACK ProAI workloadsKubernetes
0 likes · 10 min read
How ACK Pro Provisioned Control Plane Eliminates Kubernetes Control‑Plane Bottlenecks for Large‑Scale Clusters
iQIYI Technical Product Team
iQIYI Technical Product Team
Aug 6, 2026 · Artificial Intelligence

Scaling Up iQIYI’s Ad CVR Model: 10× Parameter Growth and GPU Inference Deployment

The article details how iQIYI migrated its ad conversion‑rate (CVR) model from a TensorFlow‑CPU pipeline to a TorchRec‑based PyTorch‑GPU architecture, expanding parameters tenfold while aligning offline AUC and online performance, and outlines the systematic optimizations that yielded both business metric gains and cost reductions.

CVR modelGPU inferencePerformance Optimization
0 likes · 22 min read
Scaling Up iQIYI’s Ad CVR Model: 10× Parameter Growth and GPU Inference Deployment
Architects' Tech Alliance
Architects' Tech Alliance
Aug 6, 2026 · Fundamentals

Demystifying NVIDIA GPU Core Architecture: From Basics to AI Performance

This article breaks down NVIDIA GPU fundamentals—contrasting GPU with CPU design, tracing CUDA’s evolution, detailing the hardware hierarchy from chips to SM units, explaining memory tiers, and presenting a step‑by‑step performance‑optimization methodology for AI training and inference workloads.

CUDAGPU architecturePerformance Optimization
0 likes · 22 min read
Demystifying NVIDIA GPU Core Architecture: From Basics to AI Performance
FunTester
FunTester
Aug 2, 2026 · Backend Development

Designing Effective API Caching: Strategies, Layers, and Best Practices

This guide explains API caching fundamentals, cache hit vs miss, multi‑layer cache hierarchies, common strategies such as cache‑aside and stale‑while‑revalidate, TTL tuning, event‑driven invalidation, avalanche prevention, HTTP cache directives, and key monitoring metrics to help engineers build resilient, high‑performance services.

API cachingCache invalidationCache strategies
0 likes · 10 min read
Designing Effective API Caching: Strategies, Layers, and Best Practices
ITPUB
ITPUB
Aug 2, 2026 · Operations

How to Process 10 GB of Logs in 30 Seconds with grep, sed, and awk

A senior SRE shares a step‑by‑step, performance‑focused guide on using the classic Unix trio—grep, sed, and awk—to slice, filter, and analyze massive Nginx logs, demonstrating real‑world examples, benchmark comparisons, best‑practice tips, and safety precautions for production environments.

Performance OptimizationSREawk
0 likes · 40 min read
How to Process 10 GB of Logs in 30 Seconds with grep, sed, and awk
HarmonyOS Developer Technology
HarmonyOS Developer Technology
Jul 27, 2026 · Mobile Development

ArkTS Memory Leak Detection: LocalHandle Tool Pinpoints Leaks in Minutes, Not Days

This guide introduces the LocalHandle leak detection tool for HarmonyOS ArkTS, which uses smart instrumentation to identify memory leaks at creation time by checking for active handle scopes, reducing debugging from days to minutes with 100% accuracy and direct code-line navigation in DevEco Studio.

ArkTSDebugging ToolsDevEco Studio
0 likes · 15 min read
ArkTS Memory Leak Detection: LocalHandle Tool Pinpoints Leaks in Minutes, Not Days
Java Tech Workshop
Java Tech Workshop
Jul 27, 2026 · Backend Development

Spring Boot + EasyExcel: High‑Performance, Elegant Excel Import/Export

This article explains why native Apache POI causes memory‑heavy, boilerplate Excel import/export code, introduces Alibaba's EasyExcel as a low‑memory, annotation‑driven alternative, and provides step‑by‑step Spring Boot examples for exporting, importing, custom conversion, validation, pagination, and template filling while highlighting common pitfalls and solutions.

EasyExcelExcel ExportJava
0 likes · 15 min read
Spring Boot + EasyExcel: High‑Performance, Elegant Excel Import/Export
Deepin Linux
Deepin Linux
Jul 25, 2026 · Fundamentals

Why System Calls Can Kill Performance and How to Cut Them

The article explains how frequent Linux system calls cause costly context switches, kernel checks, and cache/TLB invalidations, presents benchmark code that quantifies the overhead of getpid, open, read, and demonstrates batch I/O, caching, and algorithmic techniques to dramatically reduce those calls and boost high‑performance C++ network services.

CachingI/O batchingLinux
0 likes · 28 min read
Why System Calls Can Kill Performance and How to Cut Them
liandk
liandk
Jul 24, 2026 · Fundamentals

Why Every Project Needs Caching – Master Local and Distributed Cache Basics

The article explains the fundamental purpose of caching—placing frequently accessed data in faster storage—to dramatically reduce database load, compares local memory caches with distributed solutions like Redis, outlines their pros, cons, suitable scenarios, and presents a two‑level cache pattern plus common pitfalls such as cache penetration, breakdown, and avalanche.

CachingPerformance OptimizationRedis
0 likes · 6 min read
Why Every Project Needs Caching – Master Local and Distributed Cache Basics
IT Learning Made Simple
IT Learning Made Simple
Jul 20, 2026 · Backend Development

Key Takeaways from “Architecture Is the Future”: Scalable Web Architecture Principles

The article distills the core ideas of the book “Architecture Is the Future”, explaining why scalability is essential for modern web services and presenting eight design principles—horizontal scaling, load balancing, fault‑tolerance, data sharding, caching, asynchronous processing, monitoring, and automation—along with organizational patterns, capacity‑planning formulas, performance‑optimization steps, and high‑availability strategies.

CachingPerformance OptimizationWeb Scaling
0 likes · 11 min read
Key Takeaways from “Architecture Is the Future”: Scalable Web Architecture Principles
Machine Heart
Machine Heart
Jul 20, 2026 · Artificial Intelligence

Li Xiuhong on Cross-Cluster Heterogeneous PD Separation and Token Factory “Super Pipeline” at WAIC

The article details how 无问芯穹’s Agentic Infra strategy uses a cross‑cluster heterogeneous PD‑separation architecture (PDD) to cut first‑token latency by 51.5% and token cost by 37.5%, explains the bandwidth bottleneck of KV‑Cache transfer, introduces Decode‑side RadixCache and three‑stage handoff mechanisms, and shows a 37.5% BCR improvement that translates into roughly ten‑fold inference cost reduction.

Agentic InfraCross-ClusterLLM Inference
0 likes · 25 min read
Li Xiuhong on Cross-Cluster Heterogeneous PD Separation and Token Factory “Super Pipeline” at WAIC
ITPUB
ITPUB
Jul 17, 2026 · Backend Development

Cutting 50 M‑record Deep Paging from 10 min to 1 s – 600× Faster with ES Search‑After & Redis

This article details how a photo‑contest backend migrated from MySQL to Elasticsearch and, through three rounds of optimization—including multi‑level Redis anchor caching, recent‑anchor positioning, and a large‑interval‑plus‑small‑page‑anchor strategy—reduced arbitrary deep‑page response time from ten minutes to about one second, achieving a 600‑fold speedup while exposing remaining data‑drift challenges.

CachingDeep PaginationElasticsearch
0 likes · 14 min read
Cutting 50 M‑record Deep Paging from 10 min to 1 s – 600× Faster with ES Search‑After & Redis
Random Bulletin
Random Bulletin
Jul 14, 2026 · Backend Development

When 10 Million QPS Hits: Why Switching from Sync to Async Becomes Mandatory

The article explains how, at the ten‑million‑QPS scale, the hidden cost of thread‑bound synchronous calls—memory, scheduling overhead, and stability risks—explodes, making asynchronous architectures essential, and outlines the trade‑offs, gradual migration paths, and scenarios where async should or should not be applied.

Asynchronous ProgrammingHigh ConcurrencyPerformance Optimization
0 likes · 19 min read
When 10 Million QPS Hits: Why Switching from Sync to Async Becomes Mandatory
51CTO HarmonyOS Developer Community
51CTO HarmonyOS Developer Community
Jul 13, 2026 · Mobile Development

HarmonyOS 7 Local LLM Integration: Capabilities, Benchmarks & Engineering Guide

This article details integrating local LLMs into HarmonyOS apps, covering use cases like narrative generation and offline privacy, real-world benchmarks on Mate 60 Pro with Qwen2.5-0.5B, architecture using ArkTS and llama.cpp, performance optimizations via O3/LTO/KleidiAI, and key pitfalls like token limits and UI threading.

GGUFHarmonyOSLocal LLM
0 likes · 15 min read
HarmonyOS 7 Local LLM Integration: Capabilities, Benchmarks & Engineering Guide
Architect Chen
Architect Chen
Jul 13, 2026 · Databases

How I/O Multiplexing Gives Redis a 10× Performance Boost

Redis achieves its high speed not only because it is an in‑memory, single‑threaded database with efficient data structures, but primarily thanks to I/O multiplexing, which lets a single thread manage tens of thousands of client connections, dramatically cutting thread‑switch overhead and boosting throughput up to tenfold.

EpollHigh ConcurrencyI/O multiplexing
0 likes · 4 min read
How I/O Multiplexing Gives Redis a 10× Performance Boost
Cloud Architecture
Cloud Architecture
Jul 12, 2026 · Databases

Database Performance Optimization: 100× Speed Gains Without Changing SQL

Even without rewriting any SQL, database performance can improve up to a hundredfold by first diagnosing bottlenecks, reducing unnecessary traffic, layering read paths, optimizing indexes, tuning connection pools, and progressively evolving from a single‑node setup to read‑write separation, sharding, and distributed read models.

CachingMySQLPerformance Optimization
0 likes · 38 min read
Database Performance Optimization: 100× Speed Gains Without Changing SQL
IT Learning Made Simple
IT Learning Made Simple
Jul 12, 2026 · Game Development

From Yang Le Ge Yang to Ba Le Ge Guan: Inside the IT Architecture Behind a Viral Mini‑Game

The article breaks down the gameplay of "Ba Le Ge Guan", compares it with "Yang Le Ge Yang", and explains the three‑layer WeChat mini‑game architecture, cross‑platform rendering, performance tricks, and cloud services that make the game instantly playable, smooth, and highly engaging.

Game ArchitectureGameplay MechanicsPerformance Optimization
0 likes · 9 min read
From Yang Le Ge Yang to Ba Le Ge Guan: Inside the IT Architecture Behind a Viral Mini‑Game
Spring Full-Stack Practical Cases
Spring Full-Stack Practical Cases
Jul 11, 2026 · Backend Development

Boost Performance in Spring Boot with a Single Batch‑Processing Annotation

The article demonstrates how to create a custom @BatchProcess annotation combined with AOP to aggregate high‑frequency requests into batches, persist metadata in Redis, and process them efficiently, thereby reducing connection, I/O, and CPU overhead in distributed, high‑concurrency Spring Boot 3.5.0 applications.

AOPBatch ProcessingCustom Annotation
0 likes · 14 min read
Boost Performance in Spring Boot with a Single Batch‑Processing Annotation
Deepin Linux
Deepin Linux
Jul 11, 2026 · Backend Development

Kernel Perspective: Accelerating C++ File Transfer via System Call Optimizations

The article explains why conventional read/write loops cause excessive user‑kernel switches and data copies that inflate CPU usage and limit throughput for large or high‑frequency file transfers, and it presents zero‑copy and asynchronous I/O techniques with complete C++ examples to eliminate these bottlenecks.

C++LinuxPerformance Optimization
0 likes · 20 min read
Kernel Perspective: Accelerating C++ File Transfer via System Call Optimizations
YiSu Grain
YiSu Grain
Jul 9, 2026 · Fundamentals

How to Write Score‑Winning Answers for Architecture Case Questions

The article explains why simply listing technical terms in a software‑exam case study earns no points and provides a step‑by‑step method—identifying problems, mapping them to architectural patterns, and phrasing solutions as concrete, business‑focused sentences that score well.

CachingDistributed TransactionsPerformance Optimization
0 likes · 14 min read
How to Write Score‑Winning Answers for Architecture Case Questions
YiSu Grain
YiSu Grain
Jul 9, 2026 · Backend Development

Why Adding Cache First Is the Wrong Move for Slow Systems

The article explains that performance tuning should start with pinpointing bottlenecks using response time, throughput, concurrency and resource utilization metrics, then choose appropriate measures—caching, async processing, database tuning, horizontal scaling, and rate‑limiting—rather than blindly adding a cache.

CachingPerformance Optimizationasynchronous-processing
0 likes · 12 min read
Why Adding Cache First Is the Wrong Move for Slow Systems
IT Learning Made Simple
IT Learning Made Simple
Jul 8, 2026 · Cloud Native

From “Sheep a Sheep” to “Pull the Cupping”: Unpacking the IT Architecture and Stress‑Relief Secrets of a Hit Mini‑Game

The article breaks down the simple yet addictive gameplay of “Pull the Cupping” and explains how its three‑layer WeChat mini‑game architecture—native engine, JavaScript logic, and resource rendering—delivers cross‑platform support, performance optimization, and cloud‑backed social features, making it both a stress‑relief tool and a learning case for IT enthusiasts.

Game ArchitectureJavaScriptPerformance Optimization
0 likes · 8 min read
From “Sheep a Sheep” to “Pull the Cupping”: Unpacking the IT Architecture and Stress‑Relief Secrets of a Hit Mini‑Game
IT Learning Made Simple
IT Learning Made Simple
Jul 6, 2026 · Game Development

Dissecting the IT Architecture Behind the Viral Mini‑Game ‘Ba Le Ge Guan’

‘Ba Le Ge Guan’ blends simple, stress‑relieving gameplay with a lightweight, three‑layer architecture—native engine, JavaScript logic, and resource rendering—leveraging cross‑platform rendering, vector assets, and cloud storage to deliver instant, smooth experiences across devices, while the article compares its design to the earlier hit ‘Yang le Ge Yang’.

Game ArchitecturePerformance OptimizationWeChat Mini Game
0 likes · 9 min read
Dissecting the IT Architecture Behind the Viral Mini‑Game ‘Ba Le Ge Guan’
IT Learning Made Simple
IT Learning Made Simple
Jul 4, 2026 · Game Development

The IT Architecture Behind the Viral Mini‑Game “Cupping” and Its Stress‑Relief Appeal

The article explains how the new WeChat mini‑game “Cupping” combines a simple, stress‑relieving match‑3 mechanic with a three‑layer architecture—native engine, JavaScript game logic, and resource rendering—leveraging lightweight packaging, cross‑device rendering, performance optimizations and cloud storage to deliver instant, smooth play on any device.

Game ArchitecturePerformance OptimizationStress Relief
0 likes · 8 min read
The IT Architecture Behind the Viral Mini‑Game “Cupping” and Its Stress‑Relief Appeal
vivo Internet Technology
vivo Internet Technology
Jul 1, 2026 · Backend Development

From 10 Minutes to 1 Second: Three‑Stage Elasticsearch Deep‑Pagination Jump Optimization

This article details how a photo‑contest backend migrated from MySQL to Elasticsearch and, through three iterative optimizations—segment pre‑warming, recent‑anchor positioning with Redis ZSet, and a large‑region‑plus‑small‑page cache—reduced arbitrary deep‑page response time on 500 k records from ten minutes to under one second.

Deep PaginationElasticsearchPerformance Optimization
0 likes · 15 min read
From 10 Minutes to 1 Second: Three‑Stage Elasticsearch Deep‑Pagination Jump Optimization
Tinker Programmer
Tinker Programmer
Jun 30, 2026 · Fundamentals

Stop Using HashSet: Optimize LeetCode #3 Sliding Window from 8 ms to 2 ms

This article dissects the classic LeetCode #3 longest‑substring‑without‑repeating‑characters problem, shows why a HashSet‑based solution incurs heavy boxing overhead, and walks through three progressive optimizations—using a boolean array, index‑jumping with an int array, and refined update timing—to shrink runtime from 8 ms to about 2 ms, while highlighting common pitfalls and best‑practice guidelines.

HashSetJavaLeetCode
0 likes · 12 min read
Stop Using HashSet: Optimize LeetCode #3 Sliding Window from 8 ms to 2 ms
JD Cloud Developers
JD Cloud Developers
Jun 25, 2026 · Artificial Intelligence

JD Donates Oxygen xLLM: Open‑Source Large‑Model Inference Engine Boosts China’s AI Infrastructure

JD announced the donation of its Oxygen xLLM inference engine to the OpenAtom Open‑Source Foundation, detailing its service‑engine decoupled architecture, performance breakthroughs across e‑commerce, power and public‑safety workloads, and a roadmap to expand the open‑source AI ecosystem.

AI InfrastructureEngineering IntelligenceLarge Model Inference
0 likes · 8 min read
JD Donates Oxygen xLLM: Open‑Source Large‑Model Inference Engine Boosts China’s AI Infrastructure
JD Tech Talk
JD Tech Talk
Jun 25, 2026 · Artificial Intelligence

JD Donates Oxygen xLLM Inference Engine to OpenAtom, Boosting China’s AI Infra Ecosystem

On June 24, 2026 JD announced the donation of its Oxygen xLLM large‑model inference engine to the OpenAtom Open Source Foundation, detailing its service‑engine decoupled architecture, performance breakthroughs, heterogeneous chip support, and real‑world gains in e‑commerce, power‑grid and public‑safety applications while outlining a roadmap for broader ecosystem co‑building and standards leadership.

AI InfrastructureEngineering IntelligenceLarge Model Inference
0 likes · 7 min read
JD Donates Oxygen xLLM Inference Engine to OpenAtom, Boosting China’s AI Infra Ecosystem
JD Retail Technology
JD Retail Technology
Jun 25, 2026 · Artificial Intelligence

JD Donates Oxygen xLLM Inference Engine to OpenAtom Foundation to Accelerate Domestic AI Infra

JD announced the donation of its self‑developed Oxygen xLLM large‑model inference engine to the OpenAtom Open Source Foundation, detailing its service‑engine decoupled architecture, performance breakthroughs, multi‑chip support, and early industrial validations that aim to foster a collaborative domestic AI infrastructure ecosystem.

AI inferenceDomestic AI ecosystemEngineering Intelligence
0 likes · 8 min read
JD Donates Oxygen xLLM Inference Engine to OpenAtom Foundation to Accelerate Domestic AI Infra
Raymond Ops
Raymond Ops
Jun 25, 2026 · Operations

Linux Kernel Sysctl Tuning: Common Pitfalls and Values You Shouldn’t Change Blindly

This guide explains how to safely tune Linux kernel sysctl parameters by first identifying the problem layer, backing up current settings, applying targeted changes, and verifying effects, while highlighting common mis‑configurations, real‑world case studies, best‑practice recommendations, and monitoring strategies.

LinuxPerformance Optimizationbackup and rollback
0 likes · 18 min read
Linux Kernel Sysctl Tuning: Common Pitfalls and Values You Shouldn’t Change Blindly
Spring Full-Stack Practical Cases
Spring Full-Stack Practical Cases
Jun 24, 2026 · Backend Development

Ditch Traditional JSON Parsing: Boost Spring Boot API Performance by 30×

A four‑month investigation revealed that Jackson’s default object‑mapper consumed over 60% of CPU time during order‑submission requests, causing 900 ms latency; switching to Jackson’s streaming API reduced average response time from 912 ms to 28 ms, cut GC pauses, and increased throughput eight‑fold, while introducing readability and validation trade‑offs.

JSON parsingJacksonJava
0 likes · 8 min read
Ditch Traditional JSON Parsing: Boost Spring Boot API Performance by 30×
dbaplus Community
dbaplus Community
Jun 22, 2026 · Operations

Why Switching Linux Page Size from 4KB to 2MB Can Crash Your Performance

The article explains that blindly replacing Linux's default 4KB pages with 2MB hugepages can dramatically increase memory usage, cause cache conflicts and page‑fault latency, and ultimately degrade the performance of micro‑service workloads despite improving TLB hit rates.

HugePagesLinuxPerformance Optimization
0 likes · 19 min read
Why Switching Linux Page Size from 4KB to 2MB Can Crash Your Performance
Deepin Linux
Deepin Linux
Jun 22, 2026 · Backend Development

Memory Pool vs Object Pool: When to Choose and How to Build One from Scratch

The article explains why high‑concurrency programs suffer from memory fragmentation and system‑call overhead, compares memory pools and object pools, outlines their distinct use‑cases, provides step‑by‑step C and C++ implementations, and highlights optimization tips and common pitfalls.

C++High ConcurrencyPerformance Optimization
0 likes · 20 min read
Memory Pool vs Object Pool: When to Choose and How to Build One from Scratch
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Jun 20, 2026 · Artificial Intelligence

How I Burned $15K on Claude Code in a Month and Finally Mastered Skill Writing

After spending nearly $15,000 on Claude Code and Codex in a single month, the author discovered that most of his dozens of skills were never invoked, learned the progressive‑disclosure mechanism, rewrote skill descriptions, added verification steps, organized skills as folders with scripts and hooks, and now knows how to identify and optimize the truly useful skills.

AI agentsClaude CodePerformance Optimization
0 likes · 19 min read
How I Burned $15K on Claude Code in a Month and Finally Mastered Skill Writing
Deepin Linux
Deepin Linux
Jun 20, 2026 · Fundamentals

Why Using Pipes Can Max Out Your CPU: Hidden Costs and Fixes

Although Linux pipes avoid disk I/O and seem faster, misuse such as tiny frequent writes, mismatched read/write speeds, non‑blocking tight loops, and improper fd handling can drive a single core to 100 % CPU, but the article explains the underlying reasons and step‑by‑step optimizations to prevent it.

CPU usageIPCLinux
0 likes · 17 min read
Why Using Pipes Can Max Out Your CPU: Hidden Costs and Fixes
Spring Full-Stack Practical Cases
Spring Full-Stack Practical Cases
Jun 19, 2026 · Backend Development

Java Pooling Under High Concurrency: Resource Reuse and Performance Optimization

The article explains Java pooling techniques for high‑concurrency scenarios, introduces Apache Commons Pool 2, demonstrates how to configure dependencies, implement a PooledObjectFactory, create custom eviction policies and statistics, and shows a complete runnable example that highlights resource reuse and performance gains.

JavaPerformance OptimizationSpring Boot
0 likes · 8 min read
Java Pooling Under High Concurrency: Resource Reuse and Performance Optimization
Sohu Tech Products
Sohu Tech Products
Jun 17, 2026 · Fundamentals

Kotlin Inline Functions: More Than Just a Performance Trick

This article explains how Kotlin's inline keyword eliminates lambda object allocation and virtual calls, enables non‑local returns and reified generics, discusses inline properties and their performance benefits, and outlines scenarios where inlining can backfire, helping developers use it wisely.

AndroidKotlinPerformance Optimization
0 likes · 12 min read
Kotlin Inline Functions: More Than Just a Performance Trick
Architect's Guide
Architect's Guide
Jun 16, 2026 · Backend Development

How to Insert 300,000 Records in 13 Seconds with MyBatis and JDBC

The article compares several ways of inserting 300,000 MySQL rows—single‑row loops, an un‑batched MyBatis attempt that hits the max_allowed_packet limit, and a tuned batch strategy that commits every 1,000 rows—showing how the optimized batch reduces the runtime from hours to just 13 seconds and summarizing best‑practice tips.

Batch InsertJDBCJava
0 likes · 13 min read
How to Insert 300,000 Records in 13 Seconds with MyBatis and JDBC
HarmonyOS Developer Technology
HarmonyOS Developer Technology
Jun 15, 2026 · Mobile Development

Flutter-OH 3.35 & RNOH 0.82: Performance Gains & Architectural Shifts on HarmonyOS

The article details Flutter-OH 3.35 and RNOH 0.82 releases for HarmonyOS, covering preload rendering cutting first-frame latency 37.8%, LTPO dynamic frame rate reducing load ~20%, zero dirty-region rendering, DMA background release, BiSheng PGO, plus RNOH's parallelized pipeline, JSVM/Hermes V1 engines, Fabric-only architecture, and React 19.1.1 with DOM Node API.

Flutter 3.35Flutter-OHHarmonyOS
0 likes · 13 min read
Flutter-OH 3.35 & RNOH 0.82: Performance Gains & Architectural Shifts on HarmonyOS
HarmonyOS Developer Technology
HarmonyOS Developer Technology
Jun 15, 2026 · Industry Insights

How HarmonyOS Partners Cut Cold Starts to 1s & Boost GPU 17x at HDC2026

At HDC2026's co-construction forum, HarmonyOS partners including Amap, WPS, Tencent Video, and Meituan showcased real-world optimizations: Amap's generative UI framework boosts rendering 20%, WPS uses AI to slash flame-graph analysis from hours to minutes, Tencent Video cuts cold start to 1 second, and Meituan's QUIC tuning improves weak-network latency 27%.

AI-assisted DevelopmentGPU accelerationHDC2026
0 likes · 12 min read
How HarmonyOS Partners Cut Cold Starts to 1s & Boost GPU 17x at HDC2026
Random Bulletin
Random Bulletin
Jun 14, 2026 · Backend Development

Scaling to Millions of QPS: How Batch Consumption Beats Single-Message Processing

The article explains why single‑message consumption stalls under high QPS due to fixed per‑message overhead, and how merging pull, processing, and commit into batch operations dramatically boosts throughput while introducing trade‑offs in latency, memory, and failure handling, with practical guidelines for batch size selection and dynamic tuning.

KafkaMessage QueuePerformance Optimization
0 likes · 16 min read
Scaling to Millions of QPS: How Batch Consumption Beats Single-Message Processing
TDS Framework
TDS Framework
Jun 10, 2026 · Frontend Development

How Kuikly's TurboDisplay Achieves Instant Cross-Platform Page Loads

Tencent's Kuikly framework introduces TurboDisplay, a first-screen acceleration solution that caches rendered nodes and uses dual-thread parallel rendering to achieve instant page loads on repeat visits, reducing load times by 60-80% in production apps like QQ Games and Tencent Maps.

Kotlin MultiplatformKuiklyPerformance Optimization
0 likes · 19 min read
How Kuikly's TurboDisplay Achieves Instant Cross-Platform Page Loads
Kuaishou Tech
Kuaishou Tech
Jun 10, 2026 · Mobile Development

How Kuaishou Scaled HarmonyOS: Technical Practices Unveiled at HDC 2026

The article outlines Kuaishou's seven technical sessions at HDC 2026, detailing solutions for HarmonyOS large‑scale deployment such as startup performance, HD streaming, memory‑leak mitigation, cross‑platform framework adaptation, KMP integration, ArkUI optimization, and AI‑native enhancements.

AI integrationArkUICross-platform Development
0 likes · 9 min read
How Kuaishou Scaled HarmonyOS: Technical Practices Unveiled at HDC 2026
JD Retail Technology
JD Retail Technology
Jun 8, 2026 · Mobile Development

Accelerating Taro Native Static Layout Rendering on HarmonyOS

The article analyzes severe scroll jank on low‑end HarmonyOS devices caused by Taro Native's heavyweight card page, identifies main‑thread overload in layout phases 1, 4 and 5, proposes static node‑tree layout with custom measurement interception and font‑measurement caching, and reports a frame‑rate boost from 43 fps to 57 fps (~32.5% improvement).

CustomNodeHarmonyOSNODE_LAYOUT_RECT
0 likes · 8 min read
Accelerating Taro Native Static Layout Rendering on HarmonyOS
IT Learning Made Simple
IT Learning Made Simple
Jun 8, 2026 · R&D Management

The Essential Gear to Become a Software Architect

This guide maps the complete skill tree for aspiring software architects, detailing foundational knowledge, core competencies such as system design and performance tuning, extended expertise in cloud‑native and big‑data technologies, and a staged learning roadmap to help newcomers acquire the necessary gear.

Cloud NativePerformance OptimizationSystem Design
0 likes · 9 min read
The Essential Gear to Become a Software Architect
IT Services Circle
IT Services Circle
Jun 7, 2026 · Fundamentals

Why Switching Linux Page Size to 2 MiB Can Skyrocket Performance

The article explains how the default 4 KiB pages cause frequent TLB misses, how using 2 MiB huge pages expands a single TLB entry’s coverage by 512×, reduces page‑walk depth and page‑table overhead, and provides C++ examples for both hugetlbfs and Transparent Huge Pages.

C++Huge PagesLinux
0 likes · 7 min read
Why Switching Linux Page Size to 2 MiB Can Skyrocket Performance
Architect Practice
Architect Practice
Jun 4, 2026 · Artificial Intelligence

Why Cursor Is So Fast: Engineering Beats Model in Agent Systems (Four‑Layer Breakdown)

The article dissects Cursor's Agentic programming platform, revealing that performance gains stem from four engineering layers—architecture, inference, transport, and Agent loop—rather than merely larger models, and offers concrete optimization techniques and lessons for building fast, reliable cloud agents.

Agentic ProgrammingLinux kernel tuningPerformance Optimization
0 likes · 17 min read
Why Cursor Is So Fast: Engineering Beats Model in Agent Systems (Four‑Layer Breakdown)
360 Smart Cloud
360 Smart Cloud
Jun 2, 2026 · Databases

Valkey 9.1.0 Launches a New Era of AI‑Optimized In‑Memory Storage

Valkey 9.1.0 replaces Redis 7.2 with multi‑threaded networking, redesigned hash tables, and AI‑focused features, delivering up to 230% higher throughput, 20%+ memory savings, open BSD‑3‑Clause governance, and seamless compatibility with existing Redis ecosystems for high‑concurrency and AI workloads.

AI cachingIn-Memory DatabaseKV store
0 likes · 10 min read
Valkey 9.1.0 Launches a New Era of AI‑Optimized In‑Memory Storage
Architect Chen
Architect Chen
Jun 2, 2026 · Backend Development

Unlock 10× Faster Responses: Inside Nginx’s Caching Mechanism

The article explains how Nginx’s two‑layer caching—browser and proxy—works, why it can reduce backend load and latency, often delivering more than tenfold performance gains for read‑heavy static content, and provides detailed configuration directives such as proxy_cache_path, proxy_cache, proxy_cache_valid, and best‑practice settings to ensure cache validity and avoid cache stampede.

CachingNginxPerformance Optimization
0 likes · 5 min read
Unlock 10× Faster Responses: Inside Nginx’s Caching Mechanism
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jun 2, 2026 · Artificial Intelligence

Halving Training Time: LoongForge Full‑Stack Optimizations Boost GR00T N1.6 Throughput 2.3×

LoongForge applies system‑level optimizations—async data prefetch, fine‑grained communication‑compute overlap via a Megatron distributed optimizer, and per‑microbatch CUDA Graph scheduling—to the GR00T N1.6 Vision‑Language‑Action model, delivering up to 2.3× higher training throughput and a 56.6% reduction in overall training time on an 8×A800 cluster.

CUDA GraphDistributed TrainingGR00T N1.6
0 likes · 14 min read
Halving Training Time: LoongForge Full‑Stack Optimizations Boost GR00T N1.6 Throughput 2.3×
Woodpecker Software Testing
Woodpecker Software Testing
Jun 1, 2026 · Artificial Intelligence

Adversarial Testing Performance Optimization: Practical Strategies for Test Engineers

The article analyzes why adversarial testing is slow—highlighting redundant PGD steps, full model re‑execution, and serial verification—and presents a four‑stage optimization framework (intelligent termination, hierarchical reuse, parallel orchestration, feedback‑driven iteration) that dramatically speeds testing and enables CI/CD integration.

AI robustnessCI/CDKubernetes
0 likes · 8 min read
Adversarial Testing Performance Optimization: Practical Strategies for Test Engineers
Baidu Geek Talk
Baidu Geek Talk
Jun 1, 2026 · Cloud Computing

Cut Migration Time by 60%: How Baidu Cloud Scaled Intel Xeon 6 QAT‑Accelerated VM Live Migration

VM live migration in large cloud clusters suffers from high CPU load and long downtime; Baidu Cloud integrated Intel Xeon 6 processors with built‑in QuickAssist Technology to offload memory compression, achieving up to 60% reduction in migration duration, 20% lower CPU usage, and sub‑10 ms pause windows.

CPU OffloadCloud ComputingIntel QAT
0 likes · 10 min read
Cut Migration Time by 60%: How Baidu Cloud Scaled Intel Xeon 6 QAT‑Accelerated VM Live Migration