Read more of this story at Slashdot.
Read more of this story at Slashdot.
The C# programming language has come a remarkably long way since its early days, continuously evolving into one of the most expressive, performant, and versatile languages in modern computing. Today, with modern .NET innovations, writing high-performance enterprise applications has become smoother and more intuitive than ever.
Whether you are architecting large-scale microservices, crafting responsive web APIs, or building AI-driven solutions, keeping pace with language enhancements is essential for every developer. In this comprehensive guide, we explore the newest language capabilities, memory efficiencies, and practical patterns that will elevate your daily engineering workflow.
If you have been developing software on the Microsoft platform for a decade or more, you will recall how transformative each milestone has been. In our classic retrospective on the evolution of C# from 1.0 to 5.0, we witnessed fundamental leaps like generics, LINQ, and the async-await pattern that completely reshaped developer productivity.
Over subsequent releases, from the introduction of null-conditional operators in C# 6.0 and concise expression-bodied method syntax to modern pattern matching, the C# language team has maintained a clear focus: making code more expressive while dramatically reducing boilerplate.
Modern C# is engineered around high performance, zero-allocation memory paradigms, and cloud-native scalability. Today's language features do not simply offer syntactic sugar; they actively empower developers to write cleaner, safer code that executes with bare-metal speed across Linux containers, macOS, and Windows.
One of the most welcomed enhancements in recent C# versions is the modernization of the classic params modifier. Historically, the params keyword was strictly limited to single-dimensional arrays, meaning every method invocation with multiple arguments inevitably allocated a temporary array on the managed heap.
With expanded params collections, developers can now use params with any recognized collection type, including ReadOnlySpan<T>, Span<T>, IEnumerable<T>, List<T>, and immutable collections. This allows for clean, variadic method calls without incurring unnecessary garbage collection overhead.
Using ReadOnlySpan with params enables zero-allocation variadic methods on the stack. The compiler automatically maps arguments into a stack-allocated buffer when available, providing immediate throughput improvements for high-traffic Web API endpoints and logging utilities.
params ReadOnlySpan<T> avoids heap array creation completely, reducing GC pressure in high-frequency trading and telemetry pipelines.List<T> or custom collection types without writing separate overload wrappers for arrays.
For over two decades, C# developers synchronized concurrent threads using the lock statement alongside arbitrary reference objects (typically new object()). While familiar, this approach relied on internal synchronization blocks in the runtime header of every object, which lacked explicit synchronization intent.
Modern .NET introduces the enhanced lock object via the dedicated System.Threading.Lock type. When the C# compiler encounters a lock statement targeting a Lock instance, it generates optimized code utilizing the EnterScope() pattern instead of legacy Monitor methods.
This dedicated primitive provides clearer semantic meaning in your architecture, enables modern ref struct scoping for synchronization guards, and delivers measurable performance gains in heavily multi-threaded workloads.
myLock.EnterScope() returns a lightweight ref struct that releases the lock deterministically upon exiting the using block.
Ensuring memory safety while maintaining maximum throughput is a foundational philosophy of modern C#. Recent language versions have expanded generic constraints to support allows ref struct, enabling high-performance types like Span<T> to be used within generic abstractions for the very first time.
In addition to advanced memory mechanics, developers also enjoy subtle yet delightful daily syntax refinements. For instance, the new \e escape sequence provides a clean, standard shorthand for the ASCII escape character (0x1B), eliminating cumbersome octal or Unicode workarounds when formatting terminal outputs and ANSI color streams.
The expansion of ref struct capabilities allows developers to build high-speed parsers without sacrificing type safety. Libraries handling JSON deserialization, binary protocols, and stream parsing can now write generic algorithms that operate directly on stack memory.
allows ref struct anti-constraint permits generic interfaces and classes to accept stack-only ref structs, unlocking unprecedented performance in serialization libraries.\e escape sequence standardizes terminal formatting in CLI tools, console dashboards, and ANSI colorized loggers across platforms.
As enterprise architectures shift toward Kubernetes and serverless microservices, cold start times and memory footprints have become crucial economic factors. In cloud-native .NET applications, Native Ahead-of-Time (AOT) compilation compiles C# code directly into architecture-specific machine code without requiring a heavy JIT runtime.
Recent runtime updates have expanded Native AOT support across ASP.NET Core minimal APIs, gRPC services, and background workers. Microservices compiled with Native AOT launch in single-digit milliseconds and consume a fraction of the baseline RAM required by traditional JIT runtimes.
Moreover, developers are combining these high-speed runtimes with intelligent workflows. As shown in our tutorial on building an agentic AI workflow in C# with Microsoft AutoGen, the ecosystem provides first-class tooling for running AI orchestration and LLM integrations directly inside performant .NET services.
Through systematic performance optimization in memory allocators, vectorized SIMD instructions, and tiered compilation, .NET continues to lead industry benchmarks for web request throughput and raw computing efficiency.
Writing cutting-edge C# code is greatly enhanced by the rich developer tooling available today. In our guide on using Visual Studio for building cross-platform apps, we examined how unified IDE workflows enable building for mobile, cloud, desktop, and web from a single workstation.
Modern editions of Visual Studio and Visual Studio Code offer AI-assisted IntelliCode completions, automated refactorings for new language syntax, and integrated profiling tools that highlight memory allocations directly within your code editor.
Adopting modern language idioms not only makes your codebase more elegant and readable, but also ensures that your solutions take full advantage of runtime optimizations engineered by the Microsoft compiler teams.
The primary advantage is the combination of enhanced developer productivity and built-in performance. Features like pattern matching, record types, and expanded params allow developers to express complex business logic cleanly while minimizing heap allocations and runtime overhead.
Classic params required declaring a single-dimensional array, which always resulted in a heap allocation when passed multiple arguments. Expanded params support ReadOnlySpan, Span, List, and IEnumerable, allowing zero-allocation stack buffers and direct collection passing.
System.Threading.Lock provides dedicated synchronization semantics that the compiler and runtime optimize specifically for locking. It avoids allocating synchronization blocks in general object headers and enables clean, scope-based locking with EnterScope().
The allows ref struct anti-constraint enables generic types and methods to work with ref struct types such as Span and ReadOnlySpan. Previously, ref structs could not be used in generic parameters, which limited their reusability in high-performance generic algorithms.
Native Ahead-of-Time (AOT) compilation compiles your C# application directly into native machine code during publishing. It is ideal for cloud-native microservices, serverless functions, and containerized workloads where instant startup time and minimal memory consumption are critical.
Many syntax-level features (like pattern matching and record structs) can work with older frameworks if configured in the project file, but runtime-dependent features (such as System.Threading.Lock, Native AOT, and Span optimizations) require modern .NET runtimes.
The \e escape sequence represents the ASCII escape character (hex 0x1B, decimal 27). It provides a standard, convenient way to write ANSI escape codes for coloring and formatting text in terminal and console applications.
By providing memory-safe primitives like Span, ReadOnlySpan, ref structs, and stackalloc alongside params collections, C# enables data manipulation directly in contiguous stack memory, drastically reducing the number of objects created on the garbage-collected heap.
You can upgrade your project's Target Framework Moniker (TFM) to the latest .NET release in the .csproj file. Visual Studio and the .NET Upgrade Assistant provide automated tooling to refactor deprecated code paths into modern idioms.
Yes. With frameworks like Microsoft Semantic Kernel, AutoGen.NET, ML.NET, and ONNX Runtime bindings, C# has become a premier enterprise language for orchestrating generative AI workflows, agentic systems, and local model inference.
As we wrap up this technical overview, it is truly inspiring to see how C# continues to balance rapid modernization with robust backward compatibility. Each new language iteration provides tangible ways to write cleaner, more expressive code that simultaneously improves throughput in production environments.
I encourage you to test these new features in your day-to-day experiments and side projects. Refactor a few legacy utility classes to use params collections, try out the new Lock primitive in your background workers, and explore the benefits of Native AOT in your next microservice deployment.
What are your favorite new features in modern C#? Which language enhancements have made the biggest difference in your daily development workflow? Please share your thoughts, questions, and insights in the comments section below so we can keep the conversation going!
Thank you for reading, and happy coding!
When developers evaluate Dart for the web, they typically face a stark tradeoff:
When we built the official documentation and showcase site for BlocSignal, we knew Jaspr was the perfect foundation. But like many engineers diving into a new UI paradigm, our initial implementation took a shortcut: we used raw StatefulComponent lifecycles and manual .subscribe() callbacks to wire up our state machines.
It workedβbut it wasn't idiomatic.
In this behind-the-scenes case study, we walk through the process of dogfooding bloc_signals_jaspr across blocsignal.dev, replacing manual subscription glue with declarative consumer components, achieving 100,000 operations/sec in compiled JavaScript, and exploring the sheer developer ergonomics of Dart 3.13 primary constructors.
.subscribe() Fails at Scale
In classic Flutter or Jaspr development, when you create a state machine without framework-level consumer widgets, you might be tempted to subscribe inside initState():
// β THE ANTI-PATTERN: Manual subscription glue in StatefulComponent
class LiveVisualizerState extends State<LiveVisualizer> {
late final LiveCounterBloc _bloc;
@override
void initState() {
super.initState();
_bloc = LiveCounterBloc();
// β οΈ Flaw 1: Every state change triggers a full component setState
_bloc.state.subscribe((_) {
if (mounted) setState(() {});
});
}
@override
void dispose() {
// β οΈ Flaw 2: Manual dispose tracking
_bloc.close();
super.dispose();
}
}
While this appears harmless in a simple counter demo, it introduces three severe architectural flaws:
_bloc.add(Event()) and a local setState(), the component executes two back-to-back render passes in the exact same frame.setState() 1,000 times during the loop, creating massive JS event-loop thrashing.To solve this, we brought the full power of bloc_signals_flutter's declarative consumer components over to Jaspr in bloc_signals_jaspr.
NavigationCubit
Rather than relying on ad-hoc URL parsing scattered across components, we modeled the entire site navigation as a pure Dart state machine:
// lib/src/cubits/navigation_cubit.dart
import 'dart:js_interop';
import 'package:bloc_signals_jaspr/bloc_signals_jaspr.dart';
import 'package:web/web.dart' as web;
@JS('trackGaPageView')
external void _trackGaPageView(JSString path);
class NavigationCubit() extends CubitSignal<String> {
this : super(initialState: _resolveCurrentPath()) {
// Listen to browser history navigation
web.window.addEventListener('popstate', ((web.Event _) => _sync()).toJS);
web.window.addEventListener('hashchange', ((web.Event _) => _sync()).toJS);
_trackPageView(stateValue);
}
static String _resolveCurrentPath() {
final path = web.window.location.pathname;
final hash = web.window.location.hash.toLowerCase();
if (path.startsWith('/showcase') || hash.contains('showcase')) return '/showcase';
if (path.startsWith('/minesweeper') || hash.contains('minesweeper')) return '/minesweeper';
if (path.startsWith('/publications') || hash.contains('publications')) return '/publications';
return '/';
}
void _sync() {
final next = _resolveCurrentPath();
if (next != stateValue) {
emit(next);
_trackPageView(next);
}
}
void _trackPageView(String route) {
try {
_trackGaPageView(route.toJS);
} catch (_) {}
}
}
At the root of the application, we inject the cubit using BlocSignalProvider and build the active page using BlocSignalBuilder:
// lib/src/app.dart
class const App({super.key}) extends StatelessComponent {
@override
Component build(BuildContext context) {
return BlocSignalProvider<NavigationCubit>(
create: (_) => NavigationCubit(),
child: const _AppRouter(),
);
}
}
class const _AppRouter() extends StatelessComponent {
@override
Component build(BuildContext context) {
return BlocSignalBuilder<NavigationCubit, String>(
builder: (context, currentPath) => switch (currentPath) {
'/showcase' => const ShowcasePage(),
'/minesweeper' => const MinesweeperPage(),
'/publications' => const PublicationsPage(),
_ => const HomePage(),
},
);
}
}
Now, anywhere in the component treeβsuch as our sticky navigation headerβwe can reactively highlight active links with zero prop-drilling using context.select():
// lib/src/components/navbar.dart
final activePath = context.select<NavigationCubit, String>((c) => c.stateValue);
a(
href: '/showcase',
classes: activePath == '/showcase' ? 'nav-active' : '',
[Component.text('Showcase')],
)
BlocSignalSelector
On the blocsignal.dev homepage, the Interactive Live Visualizer demonstrates real-time state updates across multiple metrics:
state * 2.EVEN / ODD and POSITIVE / NEGATIVE / ZERO.Instead of rebuilding the entire visualizer card on every tick, each card uses BlocSignalSelector to isolate its DOM mutations:
// 1. Primary Count Selector
BlocSignalSelector<LiveCounterBloc, int, int>(
selector: (state) => state,
builder: (context, count) => span(classes: 'metric-value', [
Component.text('$count'),
]),
),
// 2. Computed 2x Doubled Selector
BlocSignalSelector<LiveCounterBloc, int, int>(
selector: (state) => state * 2,
builder: (context, doubled) => span(classes: 'metric-value', [
Component.text('$doubled'),
]),
),
// 3. Record-based Multi-Value Selector
BlocSignalSelector<LiveCounterBloc, int, ({String parity, String status})>(
selector: (state) => (
parity: state % 2 == 0 ? 'EVEN' : 'ODD',
status: state > 0 ? 'POSITIVE' : (state < 0 ? 'NEGATIVE' : 'ZERO'),
),
builder: (context, derived) => div(classes: 'status-row', [
span(classes: 'chip', [Component.text(derived.parity)]),
span(classes: 'chip', [Component.text(derived.status)]),
]),
)
Buttons dispatch events directly using context.read<LiveCounterBloc>():
button(
classes: 'btn-increment',
onClick: () => context.read<LiveCounterBloc>().add(IncrementEvent()),
[Component.text('+ 1 Increment')],
)
And background telemetry logging is captured cleanly with BlocSignalListener:
BlocSignalListener<LiveCounterBloc, int>(
listener: (context, state) {
_appendLog('β‘ TRANSITION -> State: $state [0ms Synchronous]');
},
child: visualizerMarkup,
)
Because our website is an application rather than a published library package, we can take full advantage of Dart 3.13 primary constructors and constructor shorthands.
Look at the difference in boilerplate when defining a reactive Jaspr card:
class MetricCard extends StatelessComponent {
const MetricCard({
required this.title,
required this.value,
super.key,
});
final String title;
final String value;
@override
Component build(BuildContext context) {
return div(classes: 'metric-card', [
span([Component.text(title)]),
h3([Component.text(value)]),
]);
}
}
class const MetricCard(final String title, final String value, {super.key})
extends StatelessComponent {
@override
Component build(BuildContext context) {
return div(classes: 'metric-card', [
span([Component.text(title)]),
h3([Component.text(value)]),
]);
}
}
By placing fields directly in the primary constructor parameter list, 5 lines of boilerplate collapse into a clean, single-line class header with zero repetition.
One of the biggest surprises for developers testing the live visualizer on blocsignal.dev is the built-in stress test:
Benchmark: Dispatches 1,000 synchronous transitions in a tight loop.
In traditional stream-based architectures (like classic BLoC or Rx on the web), dispatching 1,000 events allocates 1,000 StreamController events and queues 1,000 microtask hops through Dart's async runtime.
In BlocSignal, state changes propagate through a synchronous dependency graph:
Dogfooding bloc_signals_jaspr on our own production website proved that building web applications in pure Dart doesn't require choosing between developer discipline and raw performance:
| Feature | Classic Web BLoC / Rx | bloc_signals_jaspr |
|---|---|---|
| Reactivity Latency | Microtask Queue Delay (Async) | 0ms Synchronous Call Stack |
| Component Wiring | Manual subscribe / dispose
|
Declarative BlocSignalBuilder |
| DOM Rebuild Scoping | Coarse Component Rebuilds | Fine-Grained BlocSignalSelector |
| JS Web Throughput | ~2,000 β 10,000 ops/sec | ~100,000+ ops/sec |
| Syntax Overhead | Verbose Field & Constructor Maps | Dart 3.13 Primary Constructors |
You can try the interactive visualizer and play the live Minesweeper case study right now at blocsignal.dev!
All the source code is open source and visible directly in our GitHub monorepo at RandalSchwartz/BlocSignal. βοΈ
This tutorial walks through installing and setting up the Rust toolchain for vLLM on an
AWS EC2 G5g instance β Graviton2 (aarch64) with an NVIDIA T4G GPU β and getting vLLM's
Rust frontend (vllm-rs) built, running, and verified.
This paper is a follow-on to the original G5g Gemma 4 build.
Everything below was run on the box. π¦
You betcha. Since PR #40848 (merged
2026-05-21), vLLM vendors a 14-crate Rust workspace:
bench chat cmd engine-core-client llm managed-engine metrics
mock-engine parser parser/python server text tokenizer tracing
Edition 2024, resolver 3. Straight from the vendored rust/Cargo.toml:
| Crate | Version | Job |
|---|---|---|
axum |
0.8.8 | the HTTP server |
tokio |
1.47.1 | async runtime |
zeromq |
0.6.0 | talks to the Python engine |
rmp-serde / rmpv
|
1.3.1 | msgpack on the wire |
minijinja |
2.22 | chat templates |
tonic / prost
|
0.14.6 / 0.14.3 | gRPC β remember this one |
It's a drop-in replacement for the Python FastAPI server. Two artifacts get built:
vllm-rs β the axum frontend binaryvllm._rust_tool_parser β a PyO3 extension moduleThat's the headline, and it's reason enough on its own: you cannot build vLLM from source at
v0.27.2rc0 without Rust in the picture. setup.py imports it at module scope, line 21,
unguarded:
from setuptools_rust.build import build_rust
No try, no feature flag, no opt-out. Metadata generation doesn't happen without it.
And this isn't a quirk of one release. vLLM's Rust surface is 14 crates covering the HTTP
frontend, the tool parser, the tokenizer and the benchmark client, and it has been growing
since it landed. If you build inference infrastructure from source, a Rust toolchain is
becoming table stakes β so it's worth knowing how to drive it properly rather than working
around it.
Three things do get conflated, though, and they have different scopes:
| Component | Needed to build vLLM? | Needed to serve? |
|---|---|---|
setuptools_rust (Python pkg) |
yes, always | no |
cargo / rustc toolchain |
for working Rust artifacts | no |
protoc |
for vllm-rs specifically |
no |
pip install vllm need this?
Because normally pip installs it for you. pyproject.toml declares it:
[build-system]
requires = [
"cmake>=3.26.1", "ninja", "packaging>=24.2",
"setuptools>=77.0.3,<81.0.0", "setuptools-scm>=8.0",
"setuptools-rust>=1.9.0", # <- pip grabs this automatically
"torch == 2.13.0", # <- ...and this. Which is the problem.
"wheel", "jinja2",
]
Under normal build isolation, pip creates a clean env, installs that list, and builds.
You never see setuptools_rust because you never had to think about it.
But look at the torch pin. Building in isolation means pip installs torch 2.13.0 from
PyPI β and the PyPI aarch64 wheels are built for sm_80 and up. No sm_75. Which
destroys the entire reason for building from source on a T4G.
So on this box you must build against the DLAMI's own torch, and that means:
python use_existing_torch.py
pip install -e . --no-build-isolation
--no-build-isolation turns off the automatic install of everything in that requires
list. From that moment on, every build dependency is yours to supply by hand β including
setuptools_rust, which is why it turns up as a bare ModuleNotFoundError minutes into a
build that has nothing visibly to do with Rust.
So the toolchain was always required; isolation was just hiding it. Building this way means
you own the dependency list, which is the rest of this walk-through. β‘
The AWS Deep Learning ARM64 AMI ships a runtime, not a build environment. On a fresh box:
| Thing | Present? |
|---|---|
PyTorch 2.12 with sm_75
|
β |
| NVIDIA driver | β |
nvcc / CUDA toolkit |
β |
| Rust toolchain | β |
setuptools_rust |
β |
protoc |
β |
Four of those six are on you. Let's install them.
Standard rustup, nothing aarch64-specific about it:
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y
. "$HOME/.cargo/env"
stable-aarch64-unknown-linux-gnu installed - rustc 1.97.1 (8bab26f4f 2026-07-14)
Rust is installed now. Great!
Note the triple: stable-aarch64-unknown-linux-gnu. Rust's aarch64 support is a complete
non-event, which is a lovely change of pace on this hardware. β‘
python3 -m pip install setuptools_rust
Per the section above: --no-build-isolation means pip won't do this for you. Do it early
β the failure lands during metadata generation, minutes into a build, as a bare
ModuleNotFoundError: No module named 'setuptools_rust' nowhere near anything that looks
like Rust.
β οΈ Install it into the same interpreter you'll build with. On the DLAMI that's
/opt/pytorch/bin/python3, not the system python3 β they're different, and the one that
matters is whichever owns the torch you're building against.
This is the one nobody documents:
apt-get install -y protobuf-compiler
protoc --version
libprotoc 3.21.12
Why: vllm-rs depends on the vllm-server crate, vllm-server builds gRPC stubs with
tonic/prost, and prost-build shells out to protoc. Skip it and the frontend binary
does not get built β see the summary at the end for how loudly that doesn't fail.
The tool parser has no protobuf dependency, which is why it builds either way.
Not Rust, but the same class of problem, and you need it for vLLM's kernels:
# NVIDIA's **sbsa** repo β not the x86 one, easy reflex to get wrong on Arm
apt-get install -y cuda-toolkit-13-2
cd /opt/vllm-src
python tools/build_rust.py --release
β οΈ Do not omit --release. setuptools-rust builds inplace targets in debug by default,
and pip install -e . is an inplace build. The difference is not subtle:
| Artifact | Debug | Release |
|---|---|---|
_rust_tool_parser.abi3.so |
100,913,216 B | 1,009,080 B |
100x. The debug artifact is four times the size of every CUDA kernel in vLLM combined.
Timing on a g5g.xlarge (4 vCPU), cold:
real 9m1.746s
user 25m9.199s
sys 1m35.023s
501 crates. Zero warnings. Exit 0. π’
Rust's aarch64 support does not put up a fight here β which is a pleasant contrast with the
CUDA side of this box, where SM 7.5 on Graviton needs a custom arch list and a patched
kernel.
ls -la vllm/vllm-rs vllm/_rust_tool_parser.abi3.so
-rwxr-xr-x 1 root root 50039024 vllm/vllm-rs
-rwxr-xr-x 1 root root 1009080 vllm/_rust_tool_parser.abi3.so
file vllm/vllm-rs
ELF 64-bit LSB pie executable, ARM aarch64, version 1 (SYSV),
dynamically linked, interpreter /lib/ld-linux-aarch64.so.1, not stripped
vllm/vllm-rs --help
Rust frontend and managed-engine CLI for vLLM.
Commands:
frontend Run the Rust OpenAI frontend as a Python-supervised worker
serve Launch a managed Python headless engine, then run the Rust OpenAI frontend
bench Run vLLM benchmarks
render Run engine-free request rendering and preprocessing
If vllm/vllm-rs isn't there, go back to Step 3.
VLLM_USE_RUST_FRONTEND=1 vllm serve google/gemma-4-E2B-it \
--dtype float16 \
--kv-cache-dtype auto \
--max-model-len 16384 \
--gpu-memory-utilization 0.90 \
--max-num-seqs 8 \
--tensor-parallel-size 1 \
--host 0.0.0.0 --port 8000
It must be vllm serve. If you launch the module directly β
# β VLLM_USE_RUST_FRONTEND is IGNORED here
python -m vllm.entrypoints.openai.api_server --model β¦ --host 0.0.0.0 --port 8000
β the variable does nothing. No warning, no Unknown vLLM environment variable line. The
server comes up healthy and serves happily on the Python frontend, and a benchmark run
against it looks entirely normal.
The flag is read in exactly two places:
vllm/entrypoints/cli/serve.py:62 envs.VLLM_RUST_FRONTEND_PATH if envs.VLLM_USE_RUST_FRONTEND else None
vllm/entrypoints/openai/dp_supervisor.py:261 if envs.VLLM_USE_RUST_FRONTEND and envs.VLLM_RUST_FRONTEND_PATH:
api_server.py never mentions it.
Three checks. Do all three the first time.
1. The server: header:
curl -si localhost:8000/health | grep -i '^server:'
| Frontend | Response |
|---|---|
| π Python | server: uvicorn |
| π¦ Rust | (no server: header at all) |
2. The process:
pgrep -af vllm-rs
26588 /opt/vllm-src/vllm/vllm-rs frontend --listen-fd 17
--input-address ipc:///tmp/f60f3962-d45b-4bcd-9026-c0dc32736028
--output-address ipc:///tmp/5f75411d-2787-43bb-b4fc-14bf504a1cce
--engine-start-index 0 --engine-count 1 --data-parallel-size 1
3. The log prefix β (RustFrontend pid=β¦) instead of (APIServer pid=β¦):
INFO [utils.py:392] Launching Rust frontend: /opt/vllm-src/vllm/vllm-rs frontend --listen-fd 17 β¦
In two places, and they're quite different. One is a separate process; the other is a
shared object loaded inside the Python process. Here's the whole VM:
ββ EC2 g5g.4xlarge ββ Graviton2, aarch64 ββββββββββββββββββββββββββββββββββββββ
β β
β Deep Learning ARM64 AMI Β· Ubuntu 24.04 Β· NVIDIA driver 595.71.05 β
β you add > cuda-toolkit-13-2 (sbsa) Β· rustup 1.97.1 Β· protobuf-compiler β
β β
β HTTP :8000 β
β | β
β v β
β βββββββββββββββββββββββββββββββββ β
β β [RUST] vllm-rs β 50 MB aarch64 ELF, its OWN process β
β β axum 0.8.8 Β· tokio β built from the vendored rust/ workspace β
β β minijinja Β· fastokens β <- Step 5 β
β ββββββββββ¬ββββββββββββββ²βββββββββ β
β | | β
β ipc:// | ROUTER | PULL msgpack (rmp-serde / rmpv) β
β v | β
β ββββββββββ΄ββββββββββββββ΄βββββββββ β
β β [PY] vLLM supervisor β `vllm serve` opens the socket, then β
β β β hands listen-fd 17 down to vllm-rs β
β ββββββββββ¬βββββββββββββββββββββββ β
β | spawns β
β v β
β βββββββββββββββββββββββββββββββββ β
β β [PY] EngineCore β torch 2.12.0+cu132, arch list has sm_75 β
β β βββββββββββββββββββββββββββ β β
β β β [RUST] _rust_tool_parserβ β PyO3 .so LOADED INTO the Python β
β β β 1.0 MB release β β process β not a process of its own β
β β βββββββββββββββββββββββββββ β β
β ββββββββββ¬βββββββββββββββββββββββ β
β | CUDA β
β v β
β βββββββββββββββββββββββββββββββββ β
β β NVIDIA T4G Β· SM 7.5 β 15,360 MiB GDDR6 Β· 277 GB/s measured β
β β TRITON_ATTN kernels β weights 9.94 GiB Β· KV 2.95 GiB β
β βββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Two things worth pulling out of that picture:
vllm-rs is not a sidecar you point at a port. The Python side opens the listening
socket and passes the file descriptor down. It's a worker the supervisor forks and feeds._rust_tool_parser is Rust living inside Python. It's the one that always builds
(no protoc needed), which is why a broken install still leaves Rust on the box β just not
the Rust you wanted.And note where the GPU sits relative to all of this: at the bottom, behind everything. That's
the reason the benchmark below comes out the way it does.
VLLM_USE_RUST_BENCH=1 vllm bench serve β¦
Same binary, bench subcommand. Requires VLLM_RUST_FRONTEND_PATH to resolve, so it needs
the same Step 3 β Step 5 you just did.
The server came up healthy. These went by in the startup log anyway.
Gemma 4 defeats the fast tokenizer:
INFO [hf.rs:200] loading tokenizer with fastokens
WARNING [hf.rs:221] failed to load tokenizer with fastokens; falling back to
HuggingFace tokenizers
error=tokenizer error: normalizer error: unsupported normalizer type: Replace
fastokens 0.2.1 doesn't implement the Replace normalizer that Gemma 4's tokenizer.json
uses, so it falls back to the same HuggingFace tokenizers the Python path uses. Note the
fallback is graceful and correct β you just don't get the fast path on this model yet. It's
a coverage gap in a young crate, and one normalizer away from closing.
Multimodal isn't wired up for this model:
WARNING [multimodal.rs:446] multimodal model spec is not registered; disabling
image/video support model_id="google/gemma-4-E2B-it" model_type="gemma4"
Gemma 4 E2B is a vision model, and gemma4 isn't in the Rust multimodal spec table yet.
Text requests behave identically and the endpoint is healthy, so nothing in a normal check
reveals it. Also a registration gap rather than a design problem β but check it for your model
before you switch, because a healthy endpoint won't tell you.
On a T4G, no. Output token throughput, same engine config, client on the box against
localhost:
| Concurrency | π Python | π¦ Rust |
|---|---|---|
| 1 | 28.65 | 29.30 |
| 4 | 97.48 | 97.26 |
| 8 | 168.33 | 169.39 |
| 16 | 169.96 | 170.19 |
| 32 | 170.99 | 170.34 |
Median TTFT tracks just as tightly β 14305 ms against 14311 ms at concurrency 32.
That's the expected result, and worth saying plainly: decode on this card is
bandwidth-bound at a measured 277 GB/s, and the engine saturates at --max-num-seqs 8.
A frontend rewrite targets CPU-side per-request overhead. Here that overhead hides behind the
GPU, so swapping it can't move a bottleneck-limited number. If you want the Rust frontend
to buy you tokens per second on a small GPU, it won't.
One signal does appear, in median inter-token latency at high concurrency:
| Concurrency | π Python | π¦ Rust | Ξ |
|---|---|---|---|
| 16 | 38.55 | 36.18 | β6.4% |
| 32 | 38.41 | 36.23 | β5.9% |
Mean TPOT barely moves, so this is the middle of the distribution tightening rather than
everything speeding up β the shape you'd expect from a frontend scheduling streaming work
more evenly once many streams are in flight. Worth knowing if you serve at concurrency; not
worth switching for on its own. π
pip install -e . already ran
A from-source vLLM install done without the steps above succeeds, exits 0, and leaves you
with a 96 MB debug tool parser and no frontend binary. Four defaults stack up to make that
silent:
| Symptom | Cause | Fix |
|---|---|---|
No vllm/vllm-rs after a clean build |
protoc absent β vllm-server fails with code 101 |
Step 3 |
pip install exits 0 anyway |
optional=not should_require_rust_frontend() β setuptools-rust swallows it |
VLLM_REQUIRE_RUST_FRONTEND=1 |
_rust_tool_parser.abi3.so is ~96 MB |
editable β inplace β debug profile | --release |
FileNotFoundError: β¦ vllm-rs was not found |
the above, discovered at import time | Steps 3 + 5 |
Healthy server, but server: uvicorn
|
flag set on the api_server module, which never reads it |
vllm serve |
VLLM_REQUIRE_RUST_FRONTEND=1 turns the second row into a hard build failure, which is what
you want on any machine you plan to serve from.
# toolchain β you supply these by hand because the sm_75 requirement
# forces --no-build-isolation, which disables pip's automatic build deps
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y
. "$HOME/.cargo/env"
/opt/pytorch/bin/python3 -m pip install setuptools_rust # the BUILD interpreter
apt-get install -y protobuf-compiler cuda-toolkit-13-2
# build against the DLAMI's torch, not a PyPI one (PyPI aarch64 has no sm_75)
cd /path/to/vllm
python use_existing_torch.py
TORCH_CUDA_ARCH_LIST=7.5 VLLM_REQUIRE_RUST_FRONTEND=1 \
pip install -e . --no-build-isolation
# Rust artifacts, release profile (editable installs default to debug: 100x bigger)
VLLM_REQUIRE_RUST_FRONTEND=1 python tools/build_rust.py --release
# confirm
ls -la vllm/vllm-rs && vllm/vllm-rs --help
# run β `vllm serve`, NOT the api_server module
VLLM_USE_RUST_FRONTEND=1 vllm serve <model> --host 0.0.0.0 --port 8000
# verify it's really Rust
curl -si localhost:8000/health | grep -i '^server:' # Rust sends none
pgrep -af vllm-rs
Run on EC2 g5g.xlarge and g5g.4xlarge, us-east-1a, NVIDIA T4G (SM 7.5). vLLM
0.27.2rc1.dev0+g7f7a32cfe, rustc 1.97.1, setuptools-rust 1.13.0, libprotoc 3.21.12,
torch 2.12.0+cu132. Benchmarks are one run per cell for Rust and two for Python; treat the
TPOT delta as suggestive.
Robin and Mazen talk with Mike Ryan of CopilotKit about bringing AI agents to React Native apps. They break down AG-UI, generative UI, shared state, and the guardrails developers need to build useful, trustworthy mobile experiences.
Β
ShowΒ Notes
Β
ConnectΒ WithΒ Us!
Β
Sponsored by Infinite Red
Infinite Red is a premier mobile app consultancy, especially focused on Expo and React Native, located fully remote in the US. Weβre a team of 30 with highly experienced mobile app developers and have been doing this for over a decade. We are also one of the first development teams to adopt agentic coding in a way that keeps high quality standards and arenβt afraid to do things the old school way if we need to. If youβre looking for mobile app or React Native or Expo expertise for your next project, hit us up at infinite.red/radio.
In this episode, Andy sits down with Owen Fitzpatrick, a psychologist, speaker, and author of Inner Propaganda: Leading Hearts and Minds through Turbulent Times. Owen has spent close to 30 years studying how beliefs form and change, interviewing people everywhere from North Korea to Rwanda to Afghanistan. His thesis is unsettling: our brains do not simply take in facts and reach objective conclusions. They build a story we then experience as reality.
Owen and Andy work through what that means on real projects. You'll hear how a warning from a colleague can quietly harden into a conviction about a teammate, and how Bayesian reasoning gives you a way out. You'll learn Owen's five types of truth, how to tell courageous conviction from dangerous denial, and what leaders can actually make stable when they cannot promise a stable outcome. Owen also explains why pushing harder for buy-in is often the very reason people resist, and how an antifragile identity helps teams face uncertainty like AI without denial or panic.
If you're looking for a fresh way to think about belief, influence, and leading through turbulent times, this episode is for you!
You can learn more about Owen and his work at InnerPropaganda.com.
For more learning on this topic, check out:
You can chat directly with PMeLa, the podcast's AI persona, to get episode recommendations and answers to your project management and leadership questions. Visit PeopleAndProjectsPodcast.com/PMeLa to chat with her.
I know you want to be a more confident leaderβthat's why you listen to this podcast. LEAD52 is a global community of people like you who are committed to transforming their ability to lead and deliver. It's 52 weeks of leadership learning, delivered right to your inbox, taking less than 5 minutes a week. And it's all for free. Learn more and sign up at GetLEAD52.com. Thanks!
Thank you for joining me for this episode of The People and Projects Podcast!
Talent Triangle: Power Skills
Topics: Leadership, Belief, Persuasion, Influence, Uncertainty, Change Management, Growth Mindset, Psychological Reactance, Buy-In, Artificial Intelligence, Resilience, Project Management
The following music was used for this episode:
Music: The Fantastical Ferret by Tim Kulig
License (CC BY 4.0): https://filmmusic.io/standard-license
Music: Funny by Frank Schroeter
License (CC BY 4.0): https://filmmusic.io/standard-license