- Published on
AI learning journey step 5 : Notions about harness engineering
- Authors
- Name
- Ismail Tlemcani
- @Ismailtlem
In my previous post about agentic coding tools, I talked about giving agents project context. I wanted to make that more concrete in Darindex API, my FastAPI project for Moroccan real-estate data.
Most of my inspiration came from the Learn Harness Engineering course by Walking Labs. Here are the ideas I took from it and actually implemented in the API.
1. Give the agent a small entry point
The course's lesson on splitting instructions explains why putting every rule in one large file makes relevant instructions harder to find.
In Darindex, AGENTS.md gives the project overview and working rules, then directs the agent to specific documents:
| Task | Read first |
|---|---|
| Add a module or choose a layer | docs/ARCHITECTURE.md |
| Change an endpoint | docs/API.md |
| Add or edit tests | docs/testing-standards.md |
For example, the architecture document explains the route → service → repository → database structure. The API document explains response models and identifies /openapi.json as the machine-readable contract.
The agent also has to respect .ignore during routine discovery. This keeps secrets, scraped data, generated migrations, and course material out of ordinary searches, while allowing explicit inspection when needed.
2. Store project knowledge and session state in the repo
The lessons on repository knowledge and session continuity encouraged me to persist information that a new session needs.
I use two files:
PROGRESS.md: current state, active task, known issues, and next steps. It stays under 60 lines, with the current state overwritten after each task.DECISIONS.md: non-obvious decisions and their reasons.
One recorded decision explains why local tests replace developer credentials and block database connections. Another explains why shared fixtures own dependency cleanup. The next agent can read the reasoning before changing those choices.
At startup, the agent reads both files. At the end of a task, it updates progress and records blockers if the work is unfinished.
3. Keep one feature active at a time
The course's lesson on task boundaries recommends limiting work in progress so agents finish a task before expanding scope.
My AGENTS.md turns that into explicit rules: work on one feature, verify it before starting the next, and avoid unrelated refactoring.
For example, a change to district responses should stay focused on that behavior and its tests. It should not become an opportunity to reorganize the authentication service.
The workflow also requires staging only task files and committing verified changes with a descriptive message. This makes each change easier to review and resume from.
4. Make completion depend on evidence
The lesson on premature completion emphasizes execution-based evidence before declaring success.
I put the verification procedure in .agents/skills/verify/SKILL.md. It requires a baseline before implementation, checks appropriate to the change, and reporting actual results. The standard commands are:
uv run ruff check app/ tests/ main.py
uv run pytest tests/
git diff --check
The district contract test is a concrete example. It requests /cities/{city_id}/districts, checks the returned JSON, and verifies that both service calls receive the requested city ID. It runs with IDs 1 and 2, helping catch a hardcoded value.
5. Enforce test boundaries and state what they prove
The course distinguishes isolated checks from full-pipeline verification. I applied that distinction in the testing standards and verification skill.
Local route tests use shared fake services and a FastAPI test client. Before importing the app, tests/conftest.py replaces all six required settings with dummy values. An automatic fixture blocks SQLAlchemy connections:
monkeypatch.setattr(Engine, "connect", fail_connection)
monkeypatch.setattr(Engine, "raw_connection", fail_connection)
Here, fail_connection fails the test with an explanation. The shared client also clears dependency overrides in finally, even after a failed assertion.
These tests verify HTTP contracts. For database or migration changes, the documented procedure additionally requires disposable PostgreSQL with the relevant migrations applied. External provider calls must be replaced with fakes.
These are the course ideas I have applied so far: make instructions discoverable, preserve context, keep work focused, and give the agent concrete feedback about its changes.