Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161533 stories
·
33 followers

XAML.io Lets You Build .NET Apps in a Web Browser

1 Share

XAML.io is an impressive web-based software development IDE that lets you build C#/XAML .NET apps using a prompt.

The post XAML.io Lets You Build .NET Apps in a Web Browser appeared first on Thurrott.com.

Read the whole story
alvinashcraft
39 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Microsoft keeps rebuilding Windows for developers, but a new poll puts Windows at just 12%

1 Share

Apart from fixing their OS, 2026 for Microsoft was spent making Windows more attractive to developers. WSL got faster, Linux containers are now a built-in part of Windows 11, Microsoft shipped Linux-like Coreutils, and the more recent of them, Windows Developer Configurations, can turn a fresh PC into a ready-to-code machine with one command.

However, a new developer poll, spotted by Windows Latest, looks to render those efforts futile. Gergely Orosz, who runs The Pragmatic Engineer, asked his audience on X about which OS they use to build software, and got roughly 4,000 responses. macOS took a staggering 61.6%, Linux 24.2%, and Windows 12.8%

Poll on what OS developers are using to build software by Gergely Orosz
Poll on what OS developers are using to build software by Gergely Orosz. Source: Screenshot from X

Yes, Windows came in a distant third. Another poll on LinkedIn for research for the same newsletter received 6000 votes, with the numbers being 66%, 20%, and 12%, which is similar as the ones on X.

Poll numbers collected by The Pragmatic Engineer newsletter about software used by developers
Poll numbers collected by The Pragmatic Engineer newsletter about software used by developers. Source: LinkedIn

Of course, this isn’t a survey of every developer in the world. Orosz described the audience as his “bubble,” made up largely of developers at startups and Big Tech. It does not overturn the much larger surveys where Windows is still widely used.

But I feel this is an uncomfortable situation for Microsoft. Redmond is doing so much to make Windows developer-friendly, and yet the developers are still reaching for Macs and Linux boxes.

The 12% number for Windows is more of a preference than a market share

The 2025 Stack Overflow Developer Survey, with tens of thousands of respondents globally, brings us back to the real world. It found Windows to be the most widely used primary operating system among professional developers, at roughly 49.5%, compared with 32.9% for macOS.

And it being a multiple selections survey, makes it fundamentally different from Orosz’s single-choice poll. You cannot put 12.8% next to 49.5% and declare one of them wrong.

Either way, this gap between the two surveys suggests Windows is particularly weak among the highly visible startup and Big Tech crowd, even though it is still widely used across the global developer population.

Gergely Orosz about OS market share

The Statcounter chart for OS market share is misleading

The screenshot posted by Omarchy creator, which was what triggered Orosz to explain his poll, shows Statcounter’s US desktop chart with Linux near 18% and Windows above 50%.

StatCounter-os_combined-US-monthly-202508-202608
Source: Statcounter

Well, Statcounter measures page views, not people or developer machines, although its methodology says it tries to remove bot activity from a sample of billions of monthly page views.

Also, Statcounter still lists OS X and macOS as separate operating systems, years after Apple renamed OS X to macOS. Windows Latest has flagged this kind of misreading before, when a viral screenshot claimed Windows had dropped from 79% to 56% in two months. It turned out to be a classification error that Statcounter later corrected, not an exodus from Windows.

We also checked a viral news story about a spike in Linux against Cloudflare’s human versus automated traffic data and found Linux at around 4.7% for human-only North American traffic, while automated traffic pushed it as high as 26% on individual days.

Why do developers keep choosing macOS and Linux?

One word: Unix

Yes, MacBooks have their charm, but more importantly, macOS gives developers a Unix-based environment with familiar command-line tools and package managers, wrapped in a desktop that mostly stays out of the way, without bloatware and upsells. It also helps that the M-series chips are the best in class at performance and battery life.

MacBook Pro showing code for an app using Visual Studio Code
Source: Apple

There is also Xcode that has always required macOS, and it provides the only SDKs and simulators for iOS, iPadOS, watchOS, and visionOS. For anyone building Apple-platform software, choosing a Mac is the only option. And if X posts are anything to go by, Apple users are more likely to pay for apps, and hence give more revenue to the developers.

Linux is closer to where modern software runs today

Developers working with servers, containers, cloud infrastructure, and plenty of AI workloads want their development environment to resemble the Linux environments where their software eventually ships, which explains why Linux is preferred around Docker, Kubernetes, Python, and cloud infrastructure.

Microsoft’s WSL support documents mention Linux environments, containers, and GPU acceleration, while also warning that storing project files on the Windows filesystem while running Linux build tools through WSL introduces I/O overhead.

Windows is often very good at the things Microsoft controls

Windows still has enormous advantages. Visual Studio is still one of the dominant development environments, .NET runs officially across Windows, Linux, and macOS, and Windows keeps a huge enterprise footprint along with clear dominance in PC gaming.

WSL gives developers Linux access without abandoning Windows completely, and Microsoft has poured a lot of resources into Dev Drive, Windows Terminal, WinGet, and developer tooling.

The problem is that many developers do not want Windows because they need Windows. They want a Unix-like environment, and Microsoft has increasingly responded by putting Linux inside Windows.

Microsoft is practically rebuilding Windows around developers and AI

Steve Ballmer once screamed “developers, developers, developers” across a stage. It worked then, but two decades later, Microsoft is still chasing them, though the problem can’t be any more different.

WSL is no longer Microsoft’s side project

Microsoft has spent 2026 upgrading WSL with faster file access, better networking, and easier setup, describing it as foundational for running Linux workloads on Windows, work that led to WSL Containers, a built-in wslc tool for building and running Linux containers that removes the need for third-party tools like Docker Desktop.

WSL Containers in action on Windows 11
WSL Containers in action on Windows 11. Credit: Windows Latest

Funnily enough, Canonical’s VP of Engineering, Jon Seager, told The Pragmatic Engineer that Ubuntu usage inside WSL is growing faster than native Ubuntu desktop installs, and expects WSL to overtake native Ubuntu within months, driven largely by developers handed Windows laptops by their employers who then need Linux for AI and machine learning work.

Ubuntu running via Windows Subsystem for Linux
Ubuntu running via Windows Subsystem for Linux. Source: Ubuntu

Microsoft built WSL so developers would not have to leave Windows. It may be succeeding in one sense, but the fact that Microsoft had to do this at all tells us exactly what problem it was solving!

Build 2026 went full-throttle into developer tooling

At Build 2026, Microsoft pledged to make Windows 11 the OS for building AI, shipping Coreutils for Windows (a Rust-based reimplementation of GNU command-line utilities that run natively on Windows), along with Windows Developer Configurations, powered by WinGet, that set up VS Code, GitHub Copilot, WSL, and PowerShell 7 with one command.

GitHub Copilot

The same wave included Windows Development Skills for agentic native app development, an experimental Intelligent Terminal, and Microsoft Execution Containers, a policy-driven layer that lets developers control what an AI agent can touch on a device.

Project Zenith admits AI developers need a quieter Windows

Announced September 4, 2026, Project Zenith is Windows with different defaults, requiring at least 64GB of unified memory and 250GB/s of bandwidth so it can run 30-billion-parameter coding models locally.

Microsoft calls it “ready-to-code,” pinning Windows Terminal and VS Code to the taskbar and preinstalling GitHub Copilot, PowerToys, and Windows Dev Skills.

Project Zenith

Microsoft is also throwing AI-focused hardware to lure in developers. The Surface RTX Spark Dev Box and the Surface Laptop Ultra both pair NVIDIA RTX Spark with large unified-memory configurations and CUDA support, aimed squarely at local AI development. Yusuf Mehdi has pledged to spend his final year reimagining Windows for the agentic era before he leaves, and Microsoft is expected to talk more about this direction at its October Windows event, the software giant’s first major Windows event in two years.

Contrary to popular belief, Microsoft is not ignoring developers. It is adding Linux containers, improving WSL, shipping Linux-like command-line tools, building developer-specific Windows configurations, and putting AI-focused hardware to back it all. Yet an informal poll of 10,000 votes shows that Windows is still far behind macOS and Linux.

Windows is clearly capable of being a good developer platform. The harder problem is convincing developers they should want it as their development environment.

The post Microsoft keeps rebuilding Windows for developers, but a new poll puts Windows at just 12% appeared first on Windows Latest

Read the whole story
alvinashcraft
40 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Glasses, DayFrame and More!

1 Share
From: Fritz's Tech Tips and Chatter
Duration: 2:04:11
Views: 21

Feedback for DayFrame is HERE! Let's catch up...

Read the whole story
alvinashcraft
40 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

A summer of AI optimization

1 Share
A summer of AI optimization

I maintain and comaintain several open-source libraries. Some of them are widely used: ada parses URLs in Node.js, fast_float parses numbers in GCC’s standard library and in Chromium, simdjson parses JSON in Node.js, simdutf validates and transcodes Unicode in Node.js, and the Roaring bitmap libraries sit inside many database engines.

These libraries are mature. They have been optimized for years, by me and by others. For a long time, their performance was flat. Not because nobody cared, but because the remaining gains were expensive: each one required a few days of careful work, and nobody had the days.

Then, in 2026, six of them got much faster, most of it in a few weeks of summer.

To formalize my feeling, I rebuilt every commit of each library from scratch and benchmarked it on one machine (an Intel Xeon Gold 6548N). I track the speedup over time relative to August 2024. Thus the value 1.0 means no speedup. Whereas 2.0 means that the performance doubled. The lines are steps because performance only changes at a commit.

I should say that I cannot know how much AI was involved in each instance. I don’t ask how people arrived at their code. All I ask is that it be good. As for myself, I code with Claude (Opus 5), Grok and DeepSeek (V4 Pro). I was an early adopter of Grok for coding, and it got really good over time.

1. roaring (compressed bitmaps, Go)

roaring: speedup over time

The roaring library is the Go version of the Roaring index data structure. Decoding to an array got 2.5 times faster, the multi-way union FastOr got 3.1 times faster on one data set, the many-value iterator got 4.5 to 5.9 times faster, and the intersection cardinality gained 10%.

One of the contributors is an AI, actually. It is perfloop. (Disclosure: I am an advisor for perfloop.)

I did a lot of work. We also got help from Philipp Klose who declared using Claude.

2. ada (URL parsing)

ada: speedup over time

The ada library is a standard compliant URL parser. From August 2024 to July 2026, about 550 commits went in and the throughput on a corpus of 100,000 URLs stayed at 0.54 GB/s. Then, in six weeks, it went to 1.28 GB/s: 2.4 times faster, about 15 million URLs per second on one core.

Most of the optimizations were done by Yagiz Nizipli, my long-time co-author. Yagiz works at SpaceX and uses Cursor (presumably with a grok model). Abdul Rawoof Khan and Dillon Mulroy also contributed an optimization each. I worked at optimizing IP address parsing, but it won’t show in this particular benchmark.

3. fast_float (number parsing)

fast_float: speedup over time

The fast_float library parses floating-point numbers from text. It is part of GCC and most browsers. Performance was flat for fifteen months. Then, from March to July 2026, it gained 43% on one file (canada.txt, long coordinates) and 70% on another (mesh.txt, short coordinates). The optimizations should be credited to Koleman Nix and Filipe Oliveira.

4. simdjson (JSON serialization and deserialization with C++26 reflection)

simdjson: serialization and deserialization speedup over time

The simdjson library recently gained support for C++26 static reflection: you serialize and parse your own structs directly, with no glue code. Since February 2026, serialization is 1.6 times faster on twitter.json and 2.1 times faster on citm_catalog.json. Deserialization, JSON straight into a struct, gained a more modest 10% and 14% (the second panel). (The reflection code only exists since early 2026.) The number of instructions per byte fell by almost exactly the same ratio as the throughput rose: from 6.1 to 3.1 instructions per byte on citm_catalog.json serialization.

Francisco Geiman Thiesen (Microsoft) did most of the work on the serialization side while I mostly helped improve our parsing. Francisco uses Claude.

5. simdutf (Unicode validation and transcoding)

simdutf: speedup over time

The simdutf library validates and transcodes UTF-8, UTF-16 and UTF-32, and encodes and decodes base64. ASCII validation went from 83 GB/s to 160 GB/s. UTF-16 validation went from 62 GB/s to 102 GB/s. Base64 decoding gained 17%.

The work was done by Yagiz Nizipli (again) and myself.

The library got other amazing optimizations that do not show up on this benchmark by Gaspard Petit and Shreesh Adiga.

6. CRoaring (compressed bitmaps, C)

CRoaring: speedup over time

CRoaring implements Roaring bitmaps in C. On the real data sets from the repository, membership tests (contains) got 2.4 times faster, the cardinality of 64-bit bitmaps got 4.9 times faster, iterating over a 64-bit bitmap got 1.9 times faster, decoding a dense bitmap to an array got 2.2 times faster. Unions gained a more modest 13% to 16%.

The authors were Andrei Gudkov and myself.

What happened

The techniques used are all well-known. So why all these optimizations all of a sudden? Simply put, in my view, because it got cheap to try new ideas.

There is a lot of talk about the risks of AI in software. Human beings tend to be susceptible to the one-sided bet fallacy: when we see the downsides, we tend to ignore the benefits. Cars kill people, but ambulances save them.

In this instance, the benefits are concrete. Millions of people run these libraries, and this summer, they got faster.

Read the whole story
alvinashcraft
41 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

How to Port a Jekyll Blog Theme to Python: Lessons From Actually Doing It

1 Share

I've been following a tufte-jekyll styled blog for a couple of years and that led me to discover Edward Tufte's book layout.

Edward Tufte is renowned for his work on data visualization and information design, and he's a fierce advocate of high data density and for the removal of "chartjunk".

This is what tufte-css (and its many ports, including this one) brings to the web: generous whitespace, a serif reading column, and precious sidenotes for supplementary information (instead of disruptive modals).

I liked almost everything about tufe-jekyll blogs except the parts that had nothing to do with writing: a Jekyll powered Ruby version I only ever touched for this one project.

So I rewrote the whole theme in Python. Not because Jekyll is bad. It isn't. But because I wanted a toolchain I'm comfortable with. I was also curious whether I actually understood and could assimilate how a static site generator works.

Animated screenshot that displays an accessible Tufte layout template.

tufte-python is that port and this write-up acts as a guide: what actually has to happen when you move a Liquid-based Jekyll theme to a Python one, and the specific places I got it wrong before I got it right.

None of this is Jekyll-specific advice. The same pattern applies whether your target is Hugo (GoLang), Eleventy (JavaScript), or something else.

Here, you'll tinker on very focused technical points but also discover a way to break things down. If you're porting a different theme, or porting to a different language entirely, remember: the syntax changes but the shape of the challenge is the same.

Table of Contents

But Wait, Why a Static Blog?

Compared to dynamic websites, a static site has a simple publishing workflow. In this case, it consists of five steps:

  1. Write a Markdown file.

  2. Run the generator.

  3. Preview and review the result.

  4. Commit the source files.

  5. Let GitHub Actions publish the site.

This workflow is simple enough for the needs I have: occasionally publishing posts on my personal Dev blog. It keeps the content readable in a text editor and makes every change easy to review.

Like its predecessor, the actual codebase keeps Jinja2 templates, Markdown, and a YAML front matter for contents. A GitHub Actions workflow builds the site and deploys the generated _site/ directory to GitHub Pages.

Why Port a Theme Instead of Just Using It As-Is ?

There are various reasons for doing it this way.

First, maybe you want out of a toolchain you don't use anywhere else. For me that was Ruby installed on my machine for exactly and only this purpose. It was flaky enough that update my blog occasionally turned into fix my Ruby environment first.

Or maybe you already write in the target language daily, and would rather read and extend a generator you're fluent in than learn just enough of another ecosystem to (eventually) tweak a plugin file.

Or perhaps you want to understand static site generators, not just operate one. Porting forces you to read every template, every custom tag, and every build step closely enough to re-implement it. You learn and retain information differently. It's a very different level of understanding than oh! it works, and it's one of the best ways to achieve mastery.

What You'll Need

  • Basic Python: virtual environments, reading someone else's code.

  • Git and a GitHub account, since the destination for both versions is GitHub Pages.

  • Familiarity with Markdown and Git (you can even learn it as a game).

  • Basic familiarity with Jekyll's project layout: _config.yml, _layouts/, _includes/, and Liquid template syntax.

  • No prior Jinja2 experience required. It's close enough to Liquid conceptually that you'll pick it up as you go.

See the Destination First: Get the Finished Port Running

Before I get into how the port actually came together (including the parts that broke), it's worth seeing where it ends up. The theme I'm describing already exists as a ready-to-ship project. tufte-python comes with its own tutorials and you can have it running locally within minutes.

This gives you something concrete to compare against as you read the rest of this, and something to fork if you'd rather adapt an existing port than build your own from zero.

Setup

First, clone it and point it at your own repository (or fork it)

Start by creating a new, empty repository on GitHub. Give it a name such as my-blog:

git clone https://github.com/hyperphantasia/tufte-python.git my-blog
cd my-blog

Next, change the origin remote to point to your repository

git remote set-url origin <your-repository-url>

Then push the project:

git push -u origin main

Next, install the dependencies in a virtual environment

python -m venv .venv
# Uncomment to match your OS
# source .venv/bin/activate      # macOS/Linux
# .venv\Scripts\Activate.ps1     # Windows PowerShell
pip install -r requirements.txt

Now you'll want to set basic configuration values before your first build.

Open config.yml at the project root:

title: "A Quiet Corner of the Web"
author: "Your Name"
email: "you@example.com"

url: "https://yourusername.github.io"
baseurl: "/my-blog"              # "" instead, if this is a user/org page
permalink: "/articles/{year}/{slug}/"

theme: "solAArized"
options:
  mathjax: true

If you're publishing at https://yourusername.github.io/my-blog/, baseurl needs to match the repository name exactly, leading slash and no trailing slash. If you get this one wrong, every internal link and stylesheet reference on the deployed site will 404 while working fine locally (more on why in the point below).

Finally, build and preview it:

python build.py --serve --watch
Terminal of a deployed instance of tufte-python showing the localhost.

Open the address the terminal prints http://localhost:8000 (usually) and you should see the demo content that is already in content/. Leave --watch running and edit a post: the page rebuilds without you re-running anything.

One thing is worth knowing now before it costs you a confusing afternoon later: the local --serve preview ignores baseurl on purpose, so links and assets resolve from the root of your dev server instead of a subdirectory.

If you want to check the site exactly as it'll look once deployed, including the real baseurl: run python build.py --serve --production-urls instead. This is meant to preview the site using the production URL structure.

That distinction is the entire reason the works locally, breaks in production bug exists for static sites in subdirectories, and it's worth deliberately testing both modes at least once before you deploy for real.

Write your First Post

You can see the theme's features render on your own content instead of the demo's. Create content/posts/2024-06-07-hello.md:

---
title: "Hello, Margins"
date: 2024-06-07 14:30:00
categories: notes
tags: [smile, writing]
---

{% newthought 'A new thought' %} can open a section without another heading.

Here's a sidenote{% sidenote 'note-1' 'This appears in the right margin on wide screens, and behind a tap target on narrow ones.' %} to try the feature that made me want this theme in the first place.

<!--more-->

Everything past the `<!--more-->` marker stays off the homepage excerpt but shows up on the full post.

Many other visual features are available. They are discussed in details below, during implementation.

Rebuild (or let --watch pick it up), and you should see a small-caps opening phrase and a numbered note sitting in the margin next to the paragraph that references it. If both of those render, the theme's core mechanism is working end to end on your machine, good! This is the mechanic the rest of this tutorial is all about.

GitHub pages section screenshot showing the GitHub actions source to deploy correctly.

In your repository's Settings → Pages, set the source to GitHub Actions if it isn't already. The workflow bundled with the project builds and deploys automatically on every push to main. I'll walk through what that workflow is actually doing in Step 6, since GitHub Pages doesn't know what to do with a Python build script.

Ship it once you're happy with it locally:

git add config.yml content/
git commit -m "Configure site and add first post"
git push

With that running, you've got a working reference point online. Now here's how it got built.

From tufte-jekyll to tufte-python, Step by Step

To migrate a Jekyll theme to a Python build system, it's important to follow structural steps that deconstruct the existing setup.

Here, I determined six high-level steps, but that can vary depending on your task. It's very important to "own" the result in your mind first. This approach will enable you to consolidate a configuration and modernize the tooling with minimal breaks during the process.

Step 1: Inventory the Source Theme's Moving Parts

Before writing any Python, I listed every piece of Jekyll machinery the theme actually depended on. For this Liquid-heavy theme, that breaks into four categories:

Jekyll piece What it does Expected Python equivalent
_config.yml + _data/*.yml Site metadata, base URL, permalink pattern, feature toggles, structured data like social links One config.yml
_layouts/ + _includes/ Page templates and partials A templates/ directory of Jinja2 templates
_plugins/*.rb Ruby classes registering the theme's custom Liquid tags A small Python module expanding the same tag syntax
_sass/*.scss Sass partials compiled into one stylesheet at build time Plain CSS files, no compile step

I missed a fifth category on my first pass: the original theme ships two separate Rake tasks, one for scaffolding new posts and pages, and a completely different one: UploadToGithub.Rakefile for pushing the built site to a gh-pages branch by hand.

This is needed because the theme's plugins aren't in Jekyll's Pages-safe allowlist. I'd read the main Rakefile and assumed I had the whole deploy story, then wondered for some time how the original author actually got the site live.

Advice: read the whole repository root, not just the files with obvious names, before you commit to a structure.

Step 2: Collapse Scattered Config Into One File

The Jekyll version spreads settings across _config.yml (site title, URL, baseurl, permalink pattern) and one or more files under _data/: a toggle for MathJax and font loading in one file, a list of social links in another. That split follows Jekyll's own data-file conventions, but it's a complexity you don't need when you're writing your own (minimal) loader.

I consolidated all of it into a single file with clearly named sections, so anyone extending the theme later can find every setting in one place instead of three. You get something like this:

# config.yml
# --- site metadata ---
title: "A Quiet Corner of the Web"
author: "Your Name"
email: "you@example.com"

# --- URL settings ---
url: "https://yourusername.github.io"
baseurl: "/my-blog"
permalink: "/articles/{year}/{slug}/"

# --- feature toggles (previously in _data/options.yml) ---
mathjax: true
justify_text: false

# --- social links (previously in _data/social.yml) ---
social:
  - link: "github.com/yourusername"
    icon: icon-github

Step 3: Rebuild Custom Liquid Tags as Text Shortcodes

This is the part that took the longest to tinker with. It's also where most of the theme's actual personality lives. This is where you actually build the visual features: sidenotes, margin figures, and epigraphs.

How Jekyll does it

Custom Liquid tags live in _plugins/, as Ruby classes Jekyll registers with its Liquid parser. Jekyll expands them during its Liquid render pass, before handing the result to its Markdown engine.

A tag like {% sidenote "note-1" "Some aside." %} never reaches the Markdown converter as-is. It's already been swapped for HTML by the time Markdown sees the page.

Why you can't just port this 1:1 into Jinja2.

Jinja2 has its own tag system, but it's built for template-authoring logic (with loops, conditionals, and so on) not for parsing arbitrary quoted arguments out of prose sitting inside a Markdown file. And even if I'd built a Jinja2 extension for it, every existing post using the old {% sidenote ... %} syntax would need rewriting. This catch defeats the entire point of a drop-in port.

What Actually Works

Treat the tag syntax as plain text, and expand it with a preprocessing pass over the raw Markdown, before handing it to the Markdown renderer. The strategy is to mirror Jekyll's own tag-then-Markdown order exactly. A simplified version of that pass looks like this:

import re, shlex

TAG_RE = re.compile(r"\{%\s*(\w+)\s*(.*?)\s*%\}")

def split_args(raw: str) -> list[str]:
    lexer = shlex.shlex(raw, posix=True)
    lexer.whitespace_split = True
    return list(lexer)

def render_sidenote(args, resolve_img, render_md):
    note_id, text = args[0], args[1]
    text = render_md(text)
    return (f"<label for='{note_id}' class='margin-toggle sidenote-number'>"
            f"</label><input type='checkbox' id='{note_id}' "
            f"class='margin-toggle'/><span class='sidenote'>{text}</span>")

HANDLERS = {"sidenote": render_sidenote}  # All visual features are registered here

def expand_shortcodes(text: str, resolve_img, render_md) -> str:
    def dispatch(match: re.Match) -> str:
        name, raw_args = match.group(1), match.group(2)
        handler = HANDLERS.get(name)
        if handler is None:
            return match.group(0)  # leave unknown tags untouched
        return handler(split_args(raw_args), resolve_img, render_md)
    return TAG_RE.sub(dispatch, text)

The snippet above acts as a custom "search-and-replace" engine that converts shorthand tags into HTML before the final page is rendered. It uses a regular expression to scan the text for patterns like {% tag arguments %}.

The Regex (TAG_RE) is the "Scanner":

The regex is responsible for finding the tags in the big block of text. It breaks every match into two specific groups:

  • Group 1 (the name): the word immediately after {% (for example, "sidenote").

  • Group 2 (the raw arguments): everything else until the closing %} (for example, "note-1" "Some aside.").

expand_shortcodes is the "Coordinator":

This function manages the overall process. It uses re.sub to loop through the text. Every time the regex finds a match, expand_shortcodes triggers the dispatch function, which does two things:

  • It uses the name from Group 1 to look up the correct logic in the HANDLERS dictionary.

  • It passes the raw arguments from Group 2 into split_args before sending them to the parser.

split_args is the "Parser":

split_args uses the shlex library to "smart-split" the string. It recognizes quotes, so that anything inside quotation marks is kept together as a single argument. This produces a clean list where Arguments containing spaces, like a sentence inside quotes are treated as a single piece of data rather than multiple separate words (for example, ['note-1', 'Some aside.']). The final handler function can easily process that.

Render:

The last step is the actual rendering. Each tag name identified in the HANDLERS dictionary is tied to a specific Python function that knows how to return the corresponding HTML markup (for example, render_sidenote() for sidenotes).

You can have a look at the .sidenote and .margin-toggle CSS classes, to grasp an idea of how they behave visually.

Two bugs taught me why the details above matter. Both were found by throwing real old posts at the new build instead of just the demo content:

  • Quoting broke first. My first argument splitter was raw.split() on whitespace. It worked fine until I fed it a post with an apostrophe in a sidenote.

    Example: "reader's" is problematic. It split into two arguments and shift every argument after it by one. Liquid's own tag documentation actually spells out the fix: accept either single or double quotes, and allow a backslash to escape a quote inside the text. shlex in POSIX mode does exactly that in about two lines, which is a smaller fix than the bug deserved.

  • Code fences broke second. I wrote a post explaining the shortcode syntax itself, with an example wrapped in a fenced code block. This is a case of context-blindness. The regular expression is designed to find the pattern {% ... %} anywhere it appears in the document, but it doesn't know the difference between "live" code that should be executed and "example" code that is just meant to be displayed as-is to the reader. The expand_shortcodes function sees the {% and %} inside that code block and says, "Aha! A visual feature!" It then replaces the example text with the actual HTML for a sidenote and you end up seeing a broken layout where a functional feature is floating inside a code block.

    The fix is to stash fenced and inline code spans behind placeholders (like ##CODEBLOCK_1##) before running the tag regex, then restore them afterward.

The Rendered Features

The margin is not decoration. The Tufte-inspired layout remains readable thanks to the restrained typography and a generous margin set for supporting materials.

Secondary information moves into the margin instead of becoming a long interruption in the body of the article. It gives other visual elements such as notes, references, and figures a unique place to live without interrupting the main argument.

From there, porting the rest of the tags was repetitive and mechanical: same pattern, a different handler and argument count each time, the entire code is available in this file and this is how they render:

New Thought:

Tufte-Python: NewThouht example screenshot.
  • Liquid tag: {% newthought 'text' %}

Sidenote:

Tufte-Python: sidenote example screenshot.
  • Liquid tag: {% sidenote 'id' 'text' %}

    Sidenotes are numbered aside in the right margin.

Margin note:

Tufte-Python: margin note example screenshot.
  • Liquid tag: {% marginnote 'id' 'text' %}

    Margin notes are unnumbered aside in the margin.

Margin figure:

Tufte-Python: Margin figures example screenshot.
  • Liquid tag: {% marginfigure 'id' 'path' 'caption' %}

    The supporting image is confined to the margin column. Handling images isn't a big challenge, since HTML provides img tags. Positioning them correctly within the viewport is bit more tricky but was already handled well by the original SCSS.

Main column figure:

Tufte-Python: Main column figure example screenshot.
  • Liquid tag: {% maincolumn 'path' 'caption' %}

    The main image is confined to the main text column.

Full-width figure:

Tufte-Python: full width figure example screenshot.
  • Liquid tag: {% fullwidth 'path' 'caption' %}

    The full image spans on both columns.

Epigraph:

Tufte-Python: epigraph example screenshot.
  • Liquid tag: {% epigraph 'quote' 'author' 'source' %}

    This is meant for a standalone attributed quotation.

Math:

Tufte-Python: MathJax example screenshot.
  • Liquid tag: {% math %} ... {% endmath %}

    This is pure block LaTeX, rendered via MathJax.

You can also use standard markdown features, like code snippets:

Tufte-Python: code snippet example screenshot.

or tables:

Tufte-Python: table example screenshot.

These last two elements were easier to implement. Since they render in pure Markdown, it really is just about handling them directly in the CSS style sheet (for example, the Table styling section in the tufte.css file).

Step 4: Replace Compiled Sass With Swappable Plain CSS

Jekyll's Sass pipeline compiles _sass/ partials into a single stylesheet at build time, baking one fixed color palette into the output.

I didn't want a Sass-compilation dependency just to port a theme, so I stopped compiling colors into CSS.

The plain CSS stylesheet comes into two layers: structural CSS that never hardcodes a color, only references custom properties like color: var(--color-text), and one small theme file per palette that defines nothing but --color-* properties. The build copies just the selected theme's file into the output, based on a theme: key in config.yml.

Tufte python ssg animated screenshot of the available accessible themes.

Custom themes were a big improvement I wanted to implement. This turned into more than a workaround once I actually checked the numbers. I'd defaulted to Solarized first because I liked it, and only later discovered it's not optimal in terms of accessibility. That's a known, documented property: it trades some contrast for reduced eye strain.

Shipping it as the default without flagging it felt wrong for something other people might actually use to read.

Since the theme system is just swappable CSS files, the fix was adding one more file: a WCAG 2.0 AA accessible variant with the same palette adjusted to clear 4.5:1 contrast, alongside the original. That's the option this tutorial's config example points at: solAArized.

The custom-properties approach paid off again a moment later: because colors are resolved at runtime by the browser instead of baked in at build time, adding a light/dark toggle driven by prefers-color-scheme was just a small JS file to wrap. This is something a Sass-compiled single palette can't do without recompiling twice.

Step 5: Swap Filesystem-Watching for an Explicit Build Cache

jekyll serve -w bundles file-watching and incremental regeneration. Incremental involves tracking the actual state.

My first cache just tracked each post's own modification time: unchanged file, skip re-rendering. That's correct right up until you edit a shared template. I changed the post layout, rebuilt, and only two of my posts picked up the change: the ones I'd also touched that day. The others were "unchanged" by the only definition the cache knew about, so they kept their stale, pre-edit HTML in _site/.

The fix is a second, separate timestamp that isn't tied to any one document: track the newest modification time across global build inputs: templates, config.yml, and the generator's own source. If any of those is newer than the cache, force a full rebuild regardless of what any individual post's timestamp says.

import json
from pathlib import Path

CACHE_FILE = Path(".build_cache.json")

def load_cache() -> dict:
    if not CACHE_FILE.exists():
        return {"global_mtime": 0.0, "docs": {}}
    return json.loads(CACHE_FILE.read_text())

def needs_rebuild(src: Path, out: Path, cache: dict, global_stale: bool) -> bool:
    if global_stale or not out.exists():
        return True
    cached_mtime = cache["docs"].get(str(src))
    return cached_mtime is None or src.stat().st_mtime > cached_mtime

def save_cache(cache: dict, docs: dict) -> None:
    cache["docs"] = docs
    CACHE_FILE.write_text(json.dumps(cache))

The load_cache() function reads a saved JSON file that remembers when each document was last modified or it creates a fresh empty cache if the file doesn't exist yet.

The needs_rebuild() function checks whether a source file actually needs to be rebuilt by comparing its current modification time with the timestamp stored in the cache. If the file is newer than what's cached, or if the output file doesn't exist, it returns True (meaning "rebuild needed").

Finally, save_cache() updates the cache with the new build information and saves it back to the JSON file, so next time you run your build, you can skip files that haven't changed.

There's no cheap way to know which pages a shared template actually touches without re-parsing everything, so I stopped trying to be clever about it. It costs one slower build after a template edit but that's in exchange for never silently shipping a page that looks like it built successfully but didn't actually pick up the change.

Step 6: Replace Jekyll's Native GitHub Pages Build With Your Own CI

GitHub Pages knows how to build Jekyll natively. It has no idea what python build.py means, so the port needs its own CI step to build the site and hand the output to Pages:

# .github/workflows/deploy.yml
name: Build and deploy site
on:
  push:
    branches: [main]

jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install -r requirements.txt
      - run: python build.py
      - uses: actions/upload-pages-artifact@v3
        with:
          path: _site

  deploy:
    needs: build
    runs-on: ubuntu-latest
    permissions:
      pages: write
      id-token: write
    steps:
      - uses: actions/deploy-pages@v4

What happens? When you push, GitHub's servers automatically run the build job, which checks out your code, installs Python 3.12, downloads the project dependencies (from requirements.txt), runs build.py to generate the website, and then uploads the generated _site folder as an artifact.

After that succeeds, the deploy job automatically runs and takes that artifact to publish it live to GitHub Pages. Note the needs: build line. It validates the deploy step only happens after the build completes successfully, so you can't accidentally deploy a broken build.

This is the workflow the quickstart earlier in this piece relies on. Remember that in Settings → Pages, the source has to be set to GitHub Actions rather than a branch (this replaces Jekyll's built-in build step entirely). I missed that setting the first time and spent a few minutes convinced the workflow had silently failed, when it had actually succeeded and just had nowhere configured to deploy to.

Step 7: Verify Feature Parity, Not Just "It Builds"

A port that compiles cleanly isn't necessarily a correct one. Every bug I've described above passed a clean build first. Before I called it done, I tested against:

  • Real, unmodified posts from the original theme, not just demo content. This is what actually caught the quoting bug and the code-fence bug, neither of which showed up until I stopped testing against content I'd written specifically to be easy.

  • Quoting edge cases deliberately: an apostrophe inside a note, Markdown formatting inside a note, an escaped double quote.

  • Responsive behavior, since sidenotes and margin notes that tap-to-reveal on narrow screens are easy to get right on desktop and silently break on mobile versions.

Tufte python powered blog displaying a responsive state.

Responsive design is sometimes neglected and definitely not an option regarding nowadays devices diversity. Always consider it as a full and distinct user experience.

What I'd Tell Myself at the Start

Every real bug in this port came from the same root cause: testing against content I'd written to be easy, instead of content that already existed.

Don't have opinions about how the old tags should behave. The fix, every time, was the same instinct: go find the actual edge case in the old repository's documentation and code, rather than guessing at what "probably" needs to be supported.

The steps themselves generalize past this one theme: inventory the source generator's moving parts, consolidate its config, re-implement custom tags as a text-preprocessing pass instead of fighting your new template engine's syntax, swap compiled styling for something your new stack can produce without extra tooling, write your own incremental cache with an explicit escape hatch for global changes, replace whatever native deploy step you're leaving behind with your own CI, and verify against real content, not a clean build. That holds whether you're moving from Jekyll to Python, Python to Go, or anywhere else.

Thanks for reading! Feel free to contribute to tufte-python! I'm very curious about what you can come with to make this library better. More personal projects are available on my GitHub and Kaggle. You can also connect with me directly on LinkedIn as well.



Read the whole story
alvinashcraft
41 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

NOLOCK Double Counting: One Employee, Two Lunches

1 Share

NOLOCK double counting is the bug where three people stand in a lunch line and four lunches go out. Nobody new showed up. One person went through twice.

That’s this week’s episode of SQL in Sixty Seconds. Dex counts lunch tickets for Owen, June and Sam, and the total comes to four. Mia spots the problem in two seconds. Owen got back in line, and Owen’s excuse is that the second lunch is for later.

Watch It First

It’s sixty seconds, and the rest of this post makes more sense after you’ve seen Owen’s face.

Reading this in email or a feed reader where the player didn’t load? Here’s the direct link: NOLOCK Can Count the Same Row Twice, SQL in Sixty Seconds 213.

Owen, the orange robot, walks off holding a ticket marked LUNCH 4 while a purple robot with round glasses stares in surprise.

The Lunch Line Is Your Table

Swap the cafeteria for a table and the story holds up. Three employees, three rows, three IDs. A report counts them with NOLOCK, because somebody added the hint years ago to stop blocking.

SELECT COUNT(*) AS EmployeeCount
FROM dbo.Employees WITH (NOLOCK);

The answer comes back as four. There’s no fourth employee. Nothing was inserted and nothing rolled back. Owen’s row was read twice by the same scan.

How One Row Gets Two Tickets

For some scans, SQL Server reads pages in the order they sit in the file, not in key order. That can be faster when the query doesn’t need sorted output. Meanwhile, other sessions keep writing to the table.

Here’s the move from the video. The scan counts Owen’s row first. Then an update changes the column the index is sorted by. Owen’s row now belongs further along, so it moves ahead of the scan.

An animated scan has counted Owen, ID 101, once. An update lifts the same row ahead of the scan, labeled UPDATE changes sort key, same ID.

The scan keeps walking. It counts June, then Sam, then finds Owen waiting at the end. Same row, same ID, counted again.

The scan finishes past Owen's row a second time. The screen reads ACTUAL ROWS: 3, COUNT: 4, and READ TWICE BY THIS SCAN.

A page split can do the same thing. A row grows, its page runs out of room, and some rows move to a new page. If that page sits ahead of the scan, those rows get counted again. I drew both directions with animations in NOLOCK: Why It Counts Some Rows Twice and Misses Others.

What I Measured This Morning

I didn’t want to lean on a cartoon, so I tested it today on SQL Server 2025 CU8. I built a table of exactly 100,000 employees. Two sessions kept updating rows to force page splits. A third session read the table with NOLOCK, over and over, for 40 seconds.

That reader ran 413 times. 410 reads were correct. Three came back too high: 100,076, then 100,050, then 100,004. None came back short in this run.

Every one of those wrong reads still held exactly 100,000 different employee IDs. The extra rows weren’t new people. They were the same people, read twice. Employee 87076 was one of them, the Owen of my test.

All three happened in the first few seconds, while the pages were splitting hardest. That’s the part I’d remember. Double counting is rare, and it shows up when the table is busiest, which is when nobody has time to check.

Catch Owen on Your Own Table

Copy the IDs out with the same NOLOCK read, then look for any ID that shows up more than once. Run it while the table is busy. On a quiet table it returns nothing and proves nothing.

DROP TABLE IF EXISTS #Seen;

SELECT EmployeeId
INTO #Seen
FROM dbo.Employees WITH (NOLOCK);

SELECT EmployeeId, COUNT(*) AS TimesSeen
FROM #Seen
GROUP BY EmployeeId
HAVING COUNT(*) > 1;

I ran that exact script 314 times against my busy test table. Two runs came back with repeated IDs. The first one caught six employees, each seen twice. Any row it returns is an Owen: one employee, two tickets.

Getting a Count You Can Trust

If a number has to reconcile, don’t read it with NOLOCK. I covered the options in NOLOCK Can Miss Rows: The Order That Was Never Missing. Here’s the short version.

Need a rough number? Read the row counts from sys.dm_db_partition_stats. It reads metadata, not rows, so it doesn’t scan or block. Microsoft calls those counts approximate, so keep them for dashboards.

Need a number that was true at one moment? Use snapshot isolation. Readers see one consistent version of the data without blocking writers. Test it first, because the row versions need space.

Added NOLOCK to stop blocking? Fix the blocking instead. A missing index or a transaction left open will keep hurting, whatever hint the reader uses.

Owen Will Try Again

Owen’s excuse in the video is that the second lunch is for future Owen. SQL Server doesn’t have a future Owen. It has one row, counted twice, inside a total that looks perfectly normal.

I’ll admit I used to defend NOLOCK on reports that only needed a rough idea. A rough idea is fine. A number that was never true at any moment isn’t the same thing. Has a row count ever surprised you? Tell me in the comments.

NOLOCK double counting is not a new row, it is the same row read twice.

Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.

First appeared on NOLOCK Double Counting: One Employee, Two Lunches

Read the whole story
alvinashcraft
42 minutes ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories