Posts

Testing LLMs and Agents: Scoring, Consistency, and Regression

Testing LLMs is not like traditional software testing. Their stochastic nature makes consistency and reliability harder to guarantee. This week, I explore how we are testing WAAD's AI recommender using test cases, scoring rules, success rates, and regression baselines to ensure quality as we keep building new functionality.

Sell First or Build First? Plus the Client/Server Model Behind LLM Agents

Engineers and salespeople pull in opposite directions: build first, or sell first? This week I joined discovery calls with our CEO and learned why the balance matters. Plus some technical notes on client/server in LLMs and agents — why the agent is always a client of the model, and where MCP, stdio, and SSE fit in.

Journal 23 — Integration Surprises, AI Semantics, and the Question of Who I Am

This week, integration surprises exposed a gap between our AI recommendations and real ad activation. I also explore the coupling between semantic prompts and deterministic code—and the testing challenges it creates. Finally, after 25 years across development, management and startups, I ask myself: what do I do?

We Made It Out to the Lineup — Now the AI Wave Is Heading for the Shore

AI is already changing how we work in tech, from coding to automation and decision-making. I feel we have made it through the biggest part of the wave. But many professions are still standing on the shore, barely aware of what is coming. The technology is ready. The real question now is how fast society will adopt it.

Journal 20 - WordPress Still Rules, Adding a New Technology to the Stack

This week: why WordPress can still be the right choice despite more modern alternatives, and how I evaluated adding Vaadin Hilla to our Java stack. Two different decisions with the same lesson: the best technology is not always the newest one, but the one that fits the context best.

Journal 21 - Onboarding A Partner And Time To Celebrate

This week was mostly about enabling others to move fast: preparing infrastructure, access, APIs, deployment, and a secure Bedrock proxy for our Marketing partner. And after weeks of backend work, we finally reached a milestone: the first conversational UI of our recommendation product is working.

journal-week-19

This week: the human side of pushing an ERP transformation forward, dealing with resistance to change, and why friction can be a sign of progress. Plus, lessons from building an agentic app, where a small communication gap exposed an important architectural risk.

journal-week-18

Agent engineering is becoming part of every developer’s job. From MCP and agentic frameworks to harnesses and deterministic loops, I share why agents are moving from an AI specialty into everyday software engineering—and an update on our AI recommender.

Journal - Week 17

A tough but rewarding week: we paused our RAG work to focus on an agent-based recommendation engine, managed the team through another architectural change, and jumped into AWS Cognito—learning once again that AI moves faster when you know exactly where you want to go.

Journal - Week 16

This week I explore the realities of being a hands-on CTO, why we moved our recommendation agent closer to Anthropic’s capabilities, and what building production-grade software with AI taught me about focus, control, and developer quality.

Journal - Week 15

A week of tough choices: reducing scope to meet Waad’s deadline, learning how Meta, Google Ads, and LinkedIn manage advertising at scale, and facing the limits of billing and account structures. Meanwhile, an Odoo–Shopify integration reminds me that friction is often a sign that change is finally becoming real.

Journal - Week 14

This week’s reflection is about flexibility, alignment, and evolution. Business direction can change, especially in startups, and IT must be ready to adapt without losing focus or momentum. As our Marketing AI tool evolves, we are shifting from a model-based recommendation engine toward an agent-based architecture, combining frontier models with our own curated RAG dataset to create a stronger competitive advantage. At the same time, the ERP integration project is back on track, with clear milestones, better rhythm, and positive progress after the first two weeks.

Journal - Week 13

This week was all about difficult trade-offs. We reduced the MVP scope to meet an ambitious September deadline while keeping the architecture ready for future growth. We also settled on a modular monolith supported by external services, reinforced our cross-functional team approach, and I committed to turning around a struggling project. Sometimes progress isn't about flashy milestones—it's about making the right decisions that quietly set the foundation for long-term success.

Journal - Week 12

This week I reflect on two lessons: why I chose a modular monolith over microservices for an early-stage product after listening to my team, and why perseverance matters when projects don't go as planned. Architecture and leadership often have one thing in common: knowing when to change your mind.

Journal - Week 11

This week I share three topics that have been shaping my day-to-day work: using Infrastructure as Code to keep documentation synchronized with reality, experimenting with Spec-Driven Development and OpenSpec, and building an acceptance testing architecture to help a small team move faster. Different challenges, but all connected by the same goal: creating systems that scale knowledge, reduce friction, and allow teams to focus on delivering value instead of maintaining processes.

Journal - Week 10

This week reminded me that being busy and making visible progress are not always the same thing. While we successfully closed our sprint, most of the work consisted of small tasks, meetings, and alignment sessions rather than major milestones. I also continued experimenting with Spec-Driven Development and found it highly effective for individual contributors, though still challenging to scale across larger teams. On the leadership side, I reflected on how technical discussions, best practices, and business priorities help resolve disagreements. Finally, a delayed integration project taught us an important lesson: sometimes simplifying the process is the fastest path forward.

Journal - Week 9

This week I'm testing tools to keep projects in shape while building with coding agents. The problem: as code multiplies, you lose track of what was built and why. Spec-driven frameworks like OpenSpec fix this by persisting every step, not just the output. The interesting part is wiring it together—Linear, GitHub, and OpenSpec, with CLAUDE.md linking each ticket to its spec. We're still experimenting, and plenty of questions remain.

Journal - Week 8

What happens when a company loses control of half of its Facebook and Instagram presence? This week I share a real case where years of poor organization turned Meta assets into a maze of duplicate accounts, inaccessible Business Managers, and lost ownership. I also explain why I prefer adding constraints instead of policing teams, and how I use sub-milestones to keep AI projects on track without relying blindly on Agile ceremonies. A few practical lessons on organization, leadership, and project delivery from the trenches of IT.

Journal - Week 7

Choosing a cloud stack today is not only about technology — it is about trade-offs, scalability, and long-term flexibility. In this post, I explain the infrastructure decisions behind our platform: why we chose AWS over simpler PaaS alternatives, why we adopted CDK as our Infrastructure as Code solution, and how GitHub Actions fits into our CI/CD strategy. I also share the reasoning behind our serverless-first approach using ECS, Fargate, Lambda, RDS, and pgvector, plus one architectural doubt that still remains open: ECS vs EKS. If you are building modern cloud infrastructure and balancing simplicity, scalability, and maintainability, this may resonate with your own experience.

Journal - Week 6

As an IT company, or anyone building a product using agentic coding, you want to keep risk as close to 0 as possible. The product should not fail, it should be trustworthy, and at the same time it should be built as fast as possible. The same tradeoff we study in portfolio management appears here as well: lower risk usually means lower profitability, while higher profitability comes with higher risk. And I really think this framework fits the agentic coding discussion surprisingly well.

Journal - Week 5

This week’s blog post is about something we rarely discuss enough in tech leadership: The emotional side of scope negotiation. Building products is not only about architecture, estimations, or delivery plans. It is also about listening, adapting, reading people, and finding alignment between business expectations and technical reality. After weeks understanding the business, defining modules, and shaping the MVP, we finally reached the difficult conversations: What is truly essential? What can wait? Where is the real red line? Sometimes reducing scope is not failure. Sometimes it is exactly what gives a project a real chance to succeed.

Journal - Week 4

Another week building AI products made one thing clear: SaaS alone is not enough—data is the real differentiator. While interfaces and LLM features are easy to replicate, unique data is not. Building it is harder, less visible, but ultimately what creates lasting value.

Journal - Week 3

This week brought something different from the usual client work: I delivered a workshop in Logroño on Agentic Architectures — how to build apps powered by agents — to around 30 people from the local IT community. I also navigated two other challenges: helping a talented ML-focused team member get unstuck on a RAG implementation by pairing him with a seasoned app developer, and pushing a client to replace their shared-passwords Excel file with a proper password manager (slow progress, but moving in the right direction).

Journal - Week 2

A week focused on AI agentic architecture, Semantic Kernel, LangChain4j, RAGAS, and practical business challenges with Shopify and Meta. Reflections on where software engineering fits in the new AI era.

Journal - Week 1

A reflection on the first week of building an AI-driven retail product: from shaping a lean team and defining the first `showable` version, to letting go of control and embracing uncertainty in early-stage development

Deduplicating Products in Shopify

How we cleaned and deduplicated products across multiple Shopify stores during an ERP rollout—normalizing data, generating SKUs, merging catalogs safely, and protecting live revenue operations.

You build it, you run it

After being tasked with modernizing our CI/CD process, we developed a straightforward approach that significantly improved efficiency. Using a simple "Hello World" Next.js app as a reference, I realized this method could be applied across multiple technologies

D4D - Exposing the deployment as a LoadBalancer

This diagram illustrates the deployment of a Kubernetes application using AWS EC2 instances and the AWS Cloud Controller Manager (CCM). The EC2 instance, labeled as a 'K8S Node,' hosts the hello-world-lb deployment, which is exposed through a Kubernetes service configured as a LoadBalancer type. The AWS CCM interacts with the EC2 instance, enabling it to manage AWS resources like the AWS Classic Load Balancer. Security Groups are configured to allow traffic from the Load Balancer to the K8S service. The integration allows seamless connectivity between the user, AWS Load Balancer, and the Kubernetes pod running within the EC2 instance.

Hi DevOps Software Developer !

Kubernetes is not only a container orchestrator. It also serves as an infrastructure abstraction, and when embraced, it clearly defines responsibilities,

Test Post

This is a test post to validate the Hugo setup, cover image rendering, and template layouts.