AI Factory & GPU Infrastructure
GPU clusters built to actually run: Nvidia DGX and GPU nodes on Kubernetes, scheduling, storage and networking sized for real workloads - designed, deployed and validated end to end.
For 14+ years I have designed and run the infrastructure critical systems sit on - from AI Factory and GPU clusters to Kubernetes, AWS, GCP and on-prem estates, for banks, stock exchanges, industrial groups and HPC platforms. I design it, validate it, and leave it secure and documented. Now consulting independently from Alicante.
I help organisations move beyond fragile, hand-rolled systems toward infrastructure that is observable, automated, and provably secure - the kind of stack that simply does not wake people up at night.
My career has taken me from a stock exchange in Sibiu to a national security vendor, from industrial robots in the Carpathians to regulated financial institutions across Europe. The brief is always the same: make critical systems resilient, automated, and defensible - without slowing the people who depend on them.
I work end-to-end. Architecture diagrams to Ansible roles. Threat models to Grafana panels. AWS landing zones to bare-metal Kubernetes. I prefer evidence over opinion and small, durable improvements over heroic rewrites - and I hand off with runbooks and diagrams a successor could pick up cold, not tribal knowledge.
Today I operate as an independent consultant and fractional CTO serving banking, fintech, public sector and HPC clients across Europe - including AI and GPU platforms, and environments where GDPR, PCI-DSS, DORA and NIS2 are not theoretical. I speak at security conferences, mentor engineering teams, and write for openthreat.ro.
I take on a small number of long-running engagements per year. Most start as one of these; all of them end with something more resilient than they started.
GPU clusters built to actually run: Nvidia DGX and GPU nodes on Kubernetes, scheduling, storage and networking sized for real workloads - designed, deployed and validated end to end.
A structured review of your infrastructure - what's fragile, what's exposed, what won't scale - followed by benchmarking, HA and capacity validation, with an acceptance report you can hand to a board. Typically 1–3 weeks.
Landing zones, workload migrations and cost/HA re-architecture across AWS and GCP, executed with minimal downtime and a tested rollback plan.
Cluster builds, hardening and Helm-based delivery on EKS, GKE, RKE2 or bare metal - from first cluster to multi-cluster GitOps.
Senior technical leadership without a full-time hire - architecture sign-off, roadmap and vendor decisions, team mentoring, and vCISO-style security governance when you need it.
Splunk, Wazuh and Elastic deployments tuned for real detections - not just log volume - plus the automation to act on them.
Infrastructure and web-app penetration testing mapped to the frameworks your auditors actually ask about - GDPR, PCI-DSS, DORA, NIS2.
Conference talks and internal workshops on SIEM, Kubernetes and infrastructure security, drawn from real production incidents.
From university IT to senior architecture roles in cybersecurity vendors - every step shaped by the discipline of running things that cannot afford to fail.
Hands-on across the full stack - from kernel tuning to cloud landing zones - with a bias for automation, observability, and security by default.
A sample of the long-running engagements I have delivered as an independent consultant. All are under NDA and are described only at the level of scope and technology.
Long-running consultancy on a regulated AWS estate - landing zone hardening, GitOps with GitLab, Wazuh-based monitoring and serverless workloads.
Ran L3 infrastructure and security engineering across an Ansible Automation Platform estate, Kubernetes workloads, the API gateway and integration layer, and a VMware vSAN/vCenter virtualisation platform.
Modernised the Linux estate and the automation backbone across Oracle Linux and RedHat, with a Bitbucket-driven Ansible workflow and hardened integrations between core platforms.
Authored production-grade Helm charts and Kubernetes manifests for an internal platform spanning RKE2, RKE and Docker Compose footprints.
Vulnerability management, SIEM and edge protection - Nessus Pro, Wazuh, EDR and Cloudflare WAF on a Debian/Ubuntu fleet.
Designed the compute platform for a multi-tenant, GPU-accelerated environment - Nvidia DGX, Kubernetes, NSX-T and the full VMware vSphere stack.
End-to-end DevOps and observability - Ansible-driven CI/CD, Kubernetes, Splunk, Grafana/Prometheus, Nessus Pro and Veeam-backed VMware.
Cloud architecture and DevOps automation for a regulated product on AWS - VPCs, RDS, S3, SNS and Kubernetes workloads.
Small but useful - built once, shared so others can avoid the same yak-shave.
Bash script that snapshots and rotates EC2 volumes via the AWS CLI.
Python utility for stress-testing Graylog ingestion at realistic volumes.
Generates synthetic syslog streams for SIEM benchmarking and rule QA.
Web interface around the imapsync binary for safer mailbox migrations.
A small library of reusable roles for common Linux hardening and ops chores.
Curated images on Docker Hub for monitoring, log generation and ops tooling.
Continuously refreshed credentials across cloud, Kubernetes, Linux, security and emerging fields. Click any badge to verify on Credly.
→ Full credential history at credly.com/users/razvan-ioan-petrescu
Field notes from real engagements - published on openthreat.ro and presented at industry meetups.
Speaker - Security Espresso 0x0F. A working session on building maintainable, performant iptables rule sets for production Linux fleets.
securityespresso.org/0x0fCo-organiser of the OpenThreat security conference series at "1 December 1918" University of Alba Iulia - bringing practitioners and students together.
openthreat.roComputer Science training at one of Romania's oldest universities - followed by a Master's focused on advanced programming and databases.
"1 December 1918" University of Alba Iulia · Romania
2012 - 2014"1 December 1918" University of Alba Iulia · Romania
2009 - 2012The things prospective clients usually ask before an engagement starts.
An AI Factory is the compute platform your AI workloads run on: GPU nodes, scheduling, storage and networking sized to the job. I design, deploy and validate them - Nvidia DGX and GPU nodes on Kubernetes - and hand over infrastructure your team can operate, with the benchmarks to prove it performs.
A fractional CTO gives you senior technical leadership on a part-time contract: architecture sign-off, infrastructure roadmap, vendor and platform decisions, team mentoring, and security governance when you need it - without the cost of a full-time hire.
I take on a small number of long-running engagements per year and work remotely across the EU in English, Romanian, Spanish and Italian. Most start with a 1–3 week architecture or validation review and continue as hands-on delivery or ongoing advisory.
Performance benchmarking, high-availability and failover testing, capacity planning and security review - delivered as an acceptance report with findings ranked by risk and a prioritised remediation path.
Nvidia DGX and GPU infrastructure, Kubernetes (EKS, GKE, RKE2, bare metal), AWS and GCP, VMware vSphere/vSAN/Tanzu and Proxmox, with Terraform, Ansible and GitLab CI around them - plus Splunk, Wazuh and Elastic on the security side.
Yes. I work remotely with clients across the EU - including Romania and Spain - and can run the engagement in English, Romanian, Spanish or Italian. I reply within one business day.
I take on a small number of long-running engagements per year - AI Factory and GPU platforms, infrastructure architecture and validation, cloud migrations, and fractional CTO advisory for in-house teams. Tell me what you are building and what you would like to be true a quarter from now.
→ I reply within one business day.