Software engineer for web platforms and applied AI.

I’m Zendryz. I design and ship production systems - TypeScript frontends, Node APIs, and LLM infrastructure - with an emphasis on measured performance and code others can maintain.

04+
Years building and shipping
12
Systems designed end to end
02
Research tracks in progress
AreasWeb platformsLLM systems & agentsInference optimization

01 | Selected work

Work with measurable outcomes.

Open-source systems and research. Every entry links to code or to documented results - no mockups presented as shipped work.

01

Tarraco Depths

Procedural PC game prototype

Atmospheric exploration prototype with procedural layouts, spatial audio, and a custom lighting pass.

Still non-playable

Unreal Engine 5 · PC

2025 · Prototype

On request
02

Stack Composer

CLI for production scaffolds

Command-line tool that generates typed, linted, CI-ready project scaffolds instead of starter-template copy-paste.

Published on npm

Node.js · TypeScript

2025 · Shipped

03

KraftOS

Light-weight OS

Operating system with custom-made kernel capable of running Windows applications

Everything works fine

C · Assembly

2026 · Private

On request
04

GSSO

Speculative inference research

Sparse orchestration setup for LLM inference: pruned draft model plus asymmetric distillation to cut training cost.

−70% train cost · +45% throughput (local bench)

PyTorch · CUDA · Python

2026 · Research

Full archive -github.com/ZendrYz

02 - Profile

Profile.

Independent engineer, previously across web product and ML tooling work. I take on a small number of projects at a time.

I build the full system: interface, API, data layer, and - where AI is involved - the evaluation harness around the model. Most of my work sits in TypeScript and Python, and most of it is still running in production.

Current research time goes to GSSO, a speculative-decoding setup aimed at lowering inference cost.

P.01

Production over demos

Everything ships with tests, observability, and a runbook. If it only works on my machine, it doesn’t count.

P.02

Measured performance

Latency, cost per 1k requests, and bundle size are tracked from week one - not asserted in a slide deck.

P.03

Maintainable by others

Typed interfaces, documented decisions, boring technology where it fits. Handover is part of the deliverable.

Tooling I reach for
Frontend
React, Next.js, Astro, Tailwind
Backend
Node.js, PostgreSQL, Redis, Docker
AI
Python, PyTorch, LangGraph, RAG pipelines
Ops
CI/CD, error tracking, cost dashboards

03 — Services

Three engagements, clearly scoped.

Fixed scope, fixed price, weekly demos. If a project doesn’t fit one of these, I’ll say so on the first call.

S.01

Web platforms

Marketing site to logged-in product: typed frontend, API, database, auth, and billing wired correctly from day one.

Includes

  • React / Next.js / Astro builds
  • Node APIs + PostgreSQL
  • Auth, billing, roles
  • CI, previews, error tracking

01 / 03

How a project runs

  1. 01

    Scope

    Fixed written quote. What’s in, what’s out, what it costs.

  2. 02

    Prototype

    Clickable build in week one. Killed early if wrong.

  3. 03

    Build

    Weekly demos against the scope. No big-bang reveal.

  4. 04

    Handover

    Docs, runbooks, and a walkthrough. You own it after.

04 — Log

Devlog.

Build notes and measurements. Short entries, written when there’s something concrete to report.

gssoml

GSSO - Efficient Ghost Model Training

Managed to train the speculative ghost model at a fraction of the usual cost using dynamic layer pruning and asymmetric distillation. Promising throughput results.

I'd been going back and forth for weeks on how to cut down the ghost model training cost in GSSO without sacrificing speculation accuracy. The vanilla approach required training a full transformer (~7B) just to do speculative decoding - massive resource waste for a model that only verifies predictions. The fix: dynamic layer pruning + asymmetric distillation. Instead of training a full model, I train the ghost model with a dynamically pruned architecture: early layers are lightweight (low-rank attention), only the last 4-6 layers keep full resolution. Asymmetric distillation forces the ghost model to learn speculative acceptance patterns only, not the full distribution. Results: - Training cost down ~70% - Inference throughput +45% vs. standard speculative decoding - Acceptance rate dropped only 0.03 (0.92 → 0.89) - negligible given the savings Next up: implement dynamic tree drafting with this lightweight ghost model.

05 — Contact

Have a system to build? Let’s scope it.

Send a short brief: what you’re building, timeline, and budget range. You’ll get a straight answer on fit and a fixed quote - within 48 hours on working days.

Base

Valencia, Spain - working worldwide, CET timezone.

Site

© 2026 Zendryz. Back to top ↑