SKILL.md
Instance Storage & Sessions
The Pattern
Problem: You run multiple instances of something (task runners, user sessions, build jobs). Each needs isolated storage, lifecycle tracking, and its own subdirectory tree. You also have expensive initialization (loading bundles, connecting to APIs) that should happen once, not per instance.
Approach: A filesystem-backed instance store with per-instance directories, per-instance file locking, and a session factory that prepares once and creates many.
Pattern proven in production across multiple Python CLI tools and web services.
Key Design Decisions
1. Instance-rooted directory layout
All data for a single instance lives under one directory. Example layout (adapt to your needs):
~/.my-tool/
instances/{instance_id}/
instance.json — lifecycle metadata (status, timestamps)
workspace/ — code, git repo, tests
logs/ — pipeline node logs, checkpoint.json
events/ — events.jsonl, state.json, input-requests/
sidecar-data/ — sidecar service data (databases, repos)
output/ — artifacts (branch_name.txt, pr_url.txt)
This is defined as:
def get_instance_dir(instance_id: str) -> Path:
d = get_data_root() / "instances" / instance_id
d.mkdir(parents=True, exist_ok=True)
return d
# Subdirectory names used by other modules:
WORKSPACE_DIR = "workspace"
LOGS_DIR = "logs"
EVENTS_DIR = "events"
The design choice of instance-rooted (not type-rooted) matters: you can ls, tar, or rm -rf one instance directory to inspect, archive, or delete everything about that instance. Compare to the alternative where workspace files live in workspaces/{id}/ and logs in logs/{id}/ — that's harder to reason about.
2. Single env var override for all data paths
def get_data_root() -> Path:
env_dir = os.environ.get("MY_TOOL_DATA_DIR")
root = Path(env_dir) if env_dir else Path.home() / ".my-tool"
root.mkdir(parents=True, exist_ok=True)
return root
