Skip to content
View RaulMermans's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report RaulMermans

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
RaulMermans/README.md

Raúl Mermans

Applied AI & Software Engineer. I build analytical runtimes, AI agents, agent infrastructure and local-first AI tooling, and I measure whether they work.

Portfolio · LinkedIn


Selected work

Project What it is Strongest evidence
BI Notebook Lab Browser-based analytical runtime that teaches how BI semantic models compute: DAX lexer → parser → binder → evaluator, filter context, relationships, Power Query 1,024 tests · 83-case DAX conformance suite with documented divergences · 100k-row benchmark
OpsTwin Operational simulation lab for testing service-workflow changes with paired simulation, sensitivity and uncertainty ranges 247 backend + 172 frontend tests · strict typing · live deployment
DataBrief AI Bounded analytics workflow: CSV/XLSX → profiling → routed plan → sandboxed execution → grounded report 176 backend tests · every finding cites the artifact it came from
Open VS Code Agent Coding agent designed and benchmarked for a local 7B open-weight model, with explicit tools, verification and crash-safe mutation 24-task benchmark: 37.5% → 75.0% grounded success after evidence-driven orchestration
JARVIS OS Personal AI operating system built around one question, what needs my attention today?, with specialist agents, governed actions and human approval Evidence-required attention · autonomy ladder · verified execution and recovery
HALO Control Control plane for a local AI workstation: model registry, capability routing, telemetry, workload authorization Deterministic routing · metadata-only activity log · explicit mocked/real evidence classes

JARVIS OS, Open VS Code Agent and HALO Control are public architecture editions of private systems. BI Notebook Lab, OpsTwin and DataBrief AI are open source.

Progression: analytical systems (BI Notebook Lab, OpsTwin) → applied AI (DataBrief AI) → AI agents (Open VS Code Agent) → agent systems (JARVIS OS) → local AI infrastructure (HALO Control).

Open source

OpenLIT: LangGraph memory connector (open pull request). A LangGraph Store memory connector with memory CRUD/search, namespace mapping, authentication, safe content handling, docs and integration tests.

Stack

AI systems Agent orchestration (CrewAI), MCP, tool calling, memory, human-in-the-loop approval, Ollama and open-weight models
Evaluation Benchmark harnesses, conformance suites, failure taxonomies, Vitest, Pytest, Playwright
Backend TypeScript, Node.js, Fastify, Python, FastAPI, SimPy
Frontend React, Next.js, Vite, React Flow, Recharts
Data PostgreSQL, SQLite, IndexedDB, DAX and semantic modelling, Power BI concepts
Infrastructure Local inference, Vercel, GitHub Actions

Current focus

Agent systems on local and open-weight models, with an emphasis on reliability, memory, evaluation and orchestration, and on analytical runtimes that explain their own results.

Pinned Loading

  1. demand-OS demand-OS Public

    Demand forecasting and inventory-risk system that turns raw commerce data into forecasts, stockout signals, and human-reviewed reorder recommendations.

    Python

  2. OpsTwin OpsTwin Public

    Decision-support lab for testing service-workflow changes before deployment through paired discrete-event simulation, uncertainty analysis, and operational evidence.

    Python

  3. BlogAgent BlogAgent Public

    Evidence-aware editorial AI workflow that turns topics into source-grounded draft packages using query contracts, source scoring, candidate validation, claim checks, and copy-readiness gates.

    Python

  4. website-auditor website-auditor Public

    Evidence-bounded website audit workflow with deterministic scoring, public evidence capture, and bounded LLM synthesis.

    TypeScript

  5. benchmark_dashboard benchmark_dashboard Public

    Reusable competitive intelligence dashboard framework with benchmark calculations, schema validation, and executive-facing views.

    JavaScript

  6. campaign-pulse campaign-pulse Public

    Marketing intelligence dashboard for analyzing campaign performance, audience pressure, revenue, targets, and segment movement.

    TypeScript