How We Built QA11Y Labs: When AI Writes and Audits Its Own Accessible HTML
Hey everyone, Quintin here. If you've been following our journey at QA11Y Labs, you know my mission: to make the digital world genuinely accessible for everyone. And as a daily screen reader user myself, "genuine" is the operative word. No lip service, no checkbox compliance – real, usable accessibility—strictly adhering to WCAG 2.2 Level AA standards and prioritizing semantic DOM integrity over fragile ARIA overlays.
In our last post, I was pretty clear: AI, on its own, isn't going to magically solve all our accessibility problems. It's a powerful tool, sure, but it still needs human guidance, especially when it comes to the nuances of user experience. But that doesn't mean we can't leverage AI to accelerate our work and build better, more accessible systems.
The Multi-Agent Approach: Arch, Hermes, and Me
This is where our agentic AI workflow comes in. Think of it as a coordinated team: we have Arch, our orchestrator, and Hermes, our senior developer agent. And then there's me, the human accessibility expert, guiding the whole process. What we've discovered is that this multi-agent stack, when properly directed, absolutely *can* write natively accessible, visually responsive HTML, reduce ARIA bloat, and execute automated "shift-left" accessibility audits in the pipeline before deployment. It's not AI doing it alone; it's AI working with a human, specifically an accessibility-first human.
The core idea is to automate the mundane, the repetitive, and the easily verifiable, freeing up my time to focus on the complex, the experiential, and the truly impactful accessibility challenges. We're building tools that help us move faster without increasing operational risk. And we're doing it right here at QA11Y Labs.
Messy Realities: When AI Stumbles (and I Step In)
Now, it's not always a smooth, sci-fi seamless operation. There are messy realities, and that's precisely what makes this so powerful. Let me tell you about a recent moment that perfectly illustrates this.
Hermes was diligently generating some HTML for a new section of our site. The goal was to create a strict, logical Document Object Model (DOM) and semantic heading hierarchy (H1-H6) that ensures a flawless rotor navigation experience for screen reader users. I reviewed the generated code using my screen reader, as always. And then I heard it. Instead of reading out a clear heading like "Key Features," my screen reader struggled with something like "KeyFeatures."
The AI had, in its infinite wisdom, decided to concatenate two words in a heading, making it a single, unreadable blob. It wasn't a syntax error; it was an accessibility failure. A sighted user might not even notice it at first glance, but for anyone using assistive technology, it was a significant roadblock to understanding the content.
I immediately paused the process. This is where the "human accessibility expert" part of the equation becomes non-negotiable. I went into the terminal and explicitly told Hermes:
"Hermes, the heading you generated, 'KeyFeatures,' is running two words together. My screen reader pronounces it as one garbled word. I need a space between 'Key' and 'Features' so it reads as 'Key Features.' This is a critical violation of the Q-Standard for readability and semantic clarity."
And you know what? Hermes understood. It wasn't a debate or a struggle. The agent took my feedback, recognized the specific problem, and quickly applied the fix. It re-generated the HTML with the correct spacing, and when I re-audited it with my screen reader, it was perfect: "Key Features," clear as a bell.
The Q-Standard in Action
This is the Q-Standard in action. It's not just about automated checks; it's about real-world usability and my direct feedback as an accessibility specialist. We don't just trust the AI to be "good enough." We enforce our high standards, and the agents learn from every interaction.
This back-and-forth, this iterative process of AI generation, human audit, and agent correction, is how we ensure that everything we build at QA11Y Labs is natively accessible from the ground up. The AI does the heavy lifting, but I provide the critical human oversight and expertise, ensuring that the output meets the nuanced needs of all users.
So, can AI build accessible sites? Not entirely on its own. But can a well-orchestrated, multi-agent stack, guided by a human accessibility expert like me, write natively accessible HTML and self-audit its work? Absolutely. And we're proving it every day at QA11Y Labs. Stay tuned for more updates on our journey!