> INITIALIZING ROOT ENVIRONMENT...
Software Engineer specializing in high-performance C++ systems, CUDA, and applied AI. Building LLM inference engines, vector databases, and dependency-graph systems from scratch -- each verified against real references, not assumed.
> ./inspect_modules.sh
> ./view_architectures.sh
From Scratch in C++ & CUDA (Public)
A transformer inference engine written from scratch -- tokenizer, attention, KV-cache, sampling -- with no dependency on PyTorch or an existing runtime. The forward pass is verified against real HuggingFace output to a 3e-5 max logit deviation, a process that caught two real, non-crashing bugs before they shipped. Custom CUDA kernels hit 723 GFLOP/s, a 343x speedup over the CPU baseline.
Embedded Vector Database
An embedded vector database built from scratch in C++ -- a hand-implemented HNSW index, a write-ahead-logged storage engine, and scalar quantization. Benchmarked at 583µs p50 latency, 95.4% recall on SIFT -- roughly 3.5x faster than Qdrant's in-memory mode. Published to PyPI as pylattice-db.
Architectural Analytics Platform
A C++20 engine parses a repository in parallel via a std::jthread pool. A Python engine builds the dependency graph and computes real coupling, instability, and cohesion metrics. A GraphRAG-scoped engine limits AI refactoring suggestions to exactly the blast radius a change can reach. Its own CI gate blocks a pull request that pushes a module's instability past threshold. Published as a VS Code extension.
> ./verify_experience.sh
CUDA Backend Contributions -- Open Source LLM Inference Engine, 128,000+ GitHub Stars
Enabled i16 and i32 support for the GGML_OP_DUP tensor operation on CUDA, fixing a gate that was silently falling back to the CPU backend. Verified against the full backend test suite -- 16,097/16,097 tests, zero regressions -- on two Nvidia T4 GPUs. Reviewed and approved by the project's creator, Georgi Gerganov.
View PR #28897 ->Implemented a CUDA kernel for one-dimensional pooling (average and max modes), closing a gap in GPU backend coverage. Verified across 216 test cases spanning all kernel size, stride, and padding combinations on two Nvidia T4 GPUs. Merged into master by the project's creator, Georgi Gerganov.
View PR #27573 ->> ls -la /var/log/articles/
jthread, stop_token, and the deadlock I didn't see coming.
> Read TransmissionIt found a class doing 55 jobs.
> Read TransmissionScoping AI-assisted refactoring to exactly the blast radius a change can reach.
> Read TransmissionThe core algorithm behind an embedded vector database.
> Read TransmissionReal recall and latency numbers, including where it honestly loses.
> Read TransmissionThe design decisions that held up, and the ones I'd change.
> Read TransmissionTwo subtle bugs that don't crash -- they just quietly produce wrong answers.
> Read TransmissionCatching myself about to publish the flashy number instead of the honest one.
> Read TransmissionThe demo where Lattice and verbum.cpp finally talk to each other.
> Read Transmission> ./fetch_telemetry.sh
Building three production-grade systems from scratch -- an LLM inference engine, an embedded vector database, and an AI-powered architectural analytics platform -- each independently verified and benchmarked.
Deepened expertise in low-level C++ and CUDA -- custom GPU kernels, memory optimization, and systems-level performance work, culminating in a merged open-source CUDA contribution to llama.cpp.
Cyber Security and Forensics specialization. Core focus on Systems Engineering, Data Structures, Algorithms, and Low-Level Design.
University of Petroleum and Energy Studies / 2024
> ./verify_credentials.sh
Proficient in building advanced generative AI applications using RAG, agentic, and multimodal AI technologies.
Completed 15 courses covering cloud-native applications, DevOps, containers, Docker, Kubernetes, and microservices.
Mastered entry-level DevOps practices, Agile/Scrum methodologies, Python automation, and CI/CD pipelines.
> ./establish_connection.sh