velxio/backend/app/api/routes/compile.py

476 lines
18 KiB
Python
Raw Normal View History

feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
import asyncio
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
import hashlib
import logging
2026-04-26 05:46:52 +07:00
import time
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
import uuid
from typing import Any
2026-04-26 05:46:52 +07:00
from fastapi import APIRouter, Depends, HTTPException, Request
from pydantic import BaseModel
2026-04-26 05:46:52 +07:00
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
from app.core.hooks import get_current_user_id, record_compile
from app.services.arduino_cli import ArduinoCLIService
from app.services.espidf_compiler import espidf_compiler
logger = logging.getLogger(__name__)
router = APIRouter()
arduino_cli = ArduinoCLIService()
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
# ── Async compile job registry ───────────────────────────────────────────────
# In-process job dict for /compile/start + /compile/status/{job_id}. Cold ESP-IDF
# builds can take 5-7 minutes — far longer than Cloudflare's 100s edge timeout
# that hits any single HTTP request. The async path lets the client poll a
# short-lived status endpoint instead of holding one long-lived POST open.
#
# Single-instance only: if velxio ever scales to multiple FastAPI workers, this
# needs to move to Redis or the sqlite database. For now one process is fine.
COMPILE_JOBS: dict[str, dict[str, Any]] = {}
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
JOB_BY_KEY: dict[str, str] = {} # content_hash → job_id, for deduplication
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
JOB_TTL_S = 1800 # purge results 30 min after completion
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
# ── Concurrency control ──────────────────────────────────────────────────────
# Cap simultaneous ESP-IDF compiles. The VPS is modest (saw load avg 30 with
# 6 ninja processes peeling each other apart). Two parallel compiles to
# different targets are fine; concurrent compiles to the SAME target would
# corrupt the persistent build dir, so we serialize those with a per-target
# lock layered on top.
_COMPILE_SEMAPHORE = asyncio.Semaphore(2)
_TARGET_LOCKS: dict[str, asyncio.Lock] = {}
def _target_lock(board_fqbn: str) -> asyncio.Lock:
"""Lazy-initialised per-target lock so concurrent compiles to the same
board serialise. Different boards still run in parallel up to the
semaphore cap."""
lock = _TARGET_LOCKS.get(board_fqbn)
if lock is None:
lock = asyncio.Lock()
_TARGET_LOCKS[board_fqbn] = lock
return lock
def _job_key(files: list[dict[str, str]], board_fqbn: str) -> str:
"""Stable content hash of (files, board) used as deduplication key.
Excludes project_id (analytics-only different projects with identical
code should still dedup to one build). File order is normalised so the
same set of files in any order produces the same key.
"""
h = hashlib.sha256()
h.update(board_fqbn.encode())
h.update(b"\0")
for f in sorted(files, key=lambda x: x["name"]):
h.update(f["name"].encode())
h.update(b"\0")
h.update(f["content"].encode())
h.update(b"\0")
return h.hexdigest()
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
def _purge_expired_jobs() -> None:
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
"""Drop completed jobs older than JOB_TTL_S so the dict doesn't grow
forever. Also evicts the matching JOB_BY_KEY entry so the next request
with the same content schedules a fresh build instead of dedupping to
a stale job_id."""
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
now = time.time()
stale = [
jid for jid, job in COMPILE_JOBS.items()
if job.get("state") in ("done", "error")
and now - job.get("finished_at", now) > JOB_TTL_S
]
for jid in stale:
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
job = COMPILE_JOBS.pop(jid, None)
if job is not None:
key = job.get("key")
# Only remove the JOB_BY_KEY entry if it still points at this job —
# a newer job with the same key may have replaced it after this one
# finished but before TTL elapsed.
if key and JOB_BY_KEY.get(key) == jid:
JOB_BY_KEY.pop(key, None)
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
class SketchFile(BaseModel):
name: str
content: str
class CompileRequest(BaseModel):
# New multi-file API
files: list[SketchFile] | None = None
# Legacy single-file API (kept for backward compat)
code: str | None = None
board_fqbn: str = "arduino:avr:uno"
2026-04-26 05:46:52 +07:00
# Optional: associate this compile with a project for analytics
project_id: str | None = None
class CompileResponse(BaseModel):
success: bool
hex_content: str | None = None
binary_content: str | None = None # base64-encoded .bin for RP2040
binary_type: str | None = None # 'bin' or 'uf2'
has_wifi: bool = False # True when sketch uses WiFi (ESP32 only)
stdout: str
stderr: str
error: str | None = None
core_install_log: str | None = None
2026-04-26 05:46:52 +07:00
def _classify_compile_error(stderr: str, error: str | None) -> str:
"""Map raw compiler output to a stable error_kind for analytics."""
haystack = f"{error or ''}\n{stderr or ''}".lower()
if "no such file or directory" in haystack or "fatal error:" in haystack:
return "missing_library"
if "core install" in haystack or "failed to install" in haystack:
return "core_install_failed"
if "undefined reference" in haystack:
return "linker_error"
if "expected" in haystack and "before" in haystack:
return "syntax_error"
if "error:" in haystack:
return "compile_error"
return "unknown"
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
def _resolve_files(request: CompileRequest) -> list[dict[str, str]]:
"""Normalise the multi-file vs legacy single-file request bodies."""
if request.files:
return [{"name": f.name, "content": f.content} for f in request.files]
if request.code is not None:
return [{"name": "sketch.ino", "content": request.code}]
raise HTTPException(
status_code=422,
detail="Provide either 'files' or 'code' in the request body.",
)
async def _run_compile(
request: CompileRequest,
files: list[dict[str, str]],
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
progress_callback: Any = None,
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
) -> CompileResponse:
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
"""Do the actual compile (ESP-IDF for esp32:*, arduino-cli otherwise).
`progress_callback`, if provided, receives every stdout/stderr line as
cmake + ninja run. Wired into the async compile path so the live build
output is exposed via /api/compile/status/{job_id}'s `stdout` field.
AVR / RP2040 builds via arduino-cli don't surface progress yet — those
typically finish in seconds anyway.
"""
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
if request.board_fqbn.startswith("esp32:") and espidf_compiler.available:
logger.info(f"[compile] Using ESP-IDF for {request.board_fqbn}")
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
result = await espidf_compiler.compile(
files, request.board_fqbn, progress_callback=progress_callback,
)
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
return CompileResponse(
success=result["success"],
hex_content=result.get("hex_content"),
binary_content=result.get("binary_content"),
binary_type=result.get("binary_type"),
has_wifi=result.get("has_wifi", False),
stdout=result.get("stdout", ""),
stderr=result.get("stderr", ""),
error=result.get("error"),
)
# AVR, RP2040, and ESP32 fallback: use arduino-cli
core_status = await arduino_cli.ensure_core_for_board(request.board_fqbn)
core_log = core_status.get("log", "")
if core_status.get("needed") and not core_status.get("installed"):
return CompileResponse(
success=False,
stdout="",
stderr=core_log,
error=f"Failed to install required core: {core_status.get('core_id')}",
)
result = await arduino_cli.compile(files, request.board_fqbn)
return CompileResponse(
success=result["success"],
hex_content=result.get("hex_content"),
binary_content=result.get("binary_content"),
binary_type=result.get("binary_type"),
stdout=result.get("stdout", ""),
stderr=result.get("stderr", ""),
error=result.get("error"),
core_install_log=core_log if core_log else None,
)
async def _record_async_metric(
*,
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
user_id: str | None,
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
project_id: str | None,
board_fqbn: str,
success: bool,
duration_ms: int,
error_kind: str | None,
extra: dict[str, Any],
) -> None:
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
"""Forward a background-task compile metric to the registered hook.
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
Wrapper kept for the async path's signature symmetry with the sync path.
The hook owns its own DB session (the request-scoped one is gone by now)
and request=None means country/IP tagging is dropped only user_id and
timing flow through.
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
"""
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
await record_compile(
user_id=user_id,
project_id=project_id,
board_fqbn=board_fqbn,
success=success,
duration_ms=duration_ms,
error_kind=error_kind,
extra=extra,
request=None,
)
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
async def _compile_job(
job_id: str,
request: CompileRequest,
files: list[dict[str, str]],
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
user_id: str | None,
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
) -> None:
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
"""Background worker: acquire global semaphore + per-target lock, run the
compile, store result in COMPILE_JOBS.
`state=pending` while waiting on either gate; transitions to `running`
only once the actual build is about to start, so clients polling
/compile/status see an accurate snapshot of where their job is.
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
Live build output is appended to COMPILE_JOBS[job_id]['stdout_buffer']
line-by-line as cmake + ninja emit it, so /compile/status responses
stream a growing log instead of returning everything at the end.
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
"""
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
started = time.monotonic()
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
job = COMPILE_JOBS[job_id]
started_at = job["started_at"]
job_key = job.get("key")
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
# Live stdout buffer — written from a worker thread (espidf_compiler
# drain threads). dict[str].update with a single str assignment is GIL-
# protected so we don't need an explicit lock; the polling endpoint
# reads the same field.
COMPILE_JOBS[job_id]["stdout_buffer"] = ""
def on_progress_line(line: str) -> None:
# Cap buffer at 256 KB so a runaway build can't OOM the process.
# Keep the tail (most recent output) — that's what the user wants
# to see anyway.
current = COMPILE_JOBS.get(job_id)
if current is None:
return
new = (current.get("stdout_buffer", "") or "") + line
if len(new) > 262_144:
new = new[-262_144:]
current["stdout_buffer"] = new
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
try:
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
async with _COMPILE_SEMAPHORE:
async with _target_lock(request.board_fqbn):
# Job may have been purged or replaced while we were queued.
# Re-fetch and bail out if so.
if COMPILE_JOBS.get(job_id) is None:
logger.info(f"[compile] job {job_id} purged before run; skipping")
return
COMPILE_JOBS[job_id]["state"] = "running"
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
response = await _run_compile(
request, files, progress_callback=on_progress_line,
)
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
COMPILE_JOBS[job_id] = {
"state": "done",
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
"started_at": started_at,
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
"finished_at": time.time(),
"result": response.model_dump(),
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
"key": job_key,
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
# Preserve the streamed buffer post-completion so a late poll
# still has access to the live log (clients usually display
# result.stdout once state=done, but having both costs nothing).
"stdout_buffer": COMPILE_JOBS.get(job_id, {}).get("stdout_buffer", ""),
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
}
error_kind = (
None if response.success
else _classify_compile_error(response.stderr, response.error)
)
await _record_async_metric(
user_id=user_id,
project_id=request.project_id,
board_fqbn=request.board_fqbn,
success=response.success,
duration_ms=int((time.monotonic() - started) * 1000),
error_kind=error_kind,
extra={"file_count": len(files), "has_wifi": response.has_wifi, "async": True},
)
except Exception as exc:
logger.exception(f"[compile] async job {job_id} failed")
COMPILE_JOBS[job_id] = {
"state": "error",
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
"started_at": started_at,
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
"finished_at": time.time(),
"error": str(exc)[:500],
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
"key": job_key,
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
"stdout_buffer": COMPILE_JOBS.get(job_id, {}).get("stdout_buffer", ""),
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
}
await _record_async_metric(
user_id=user_id,
project_id=request.project_id,
board_fqbn=request.board_fqbn,
success=False,
duration_ms=int((time.monotonic() - started) * 1000),
error_kind="exception",
extra={"file_count": len(files), "exception": str(exc)[:200], "async": True},
)
@router.post("/", response_model=CompileResponse)
2026-04-26 05:46:52 +07:00
async def compile_sketch(
request: CompileRequest,
http_request: Request,
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
user_id: str | None = Depends(get_current_user_id),
2026-04-26 05:46:52 +07:00
):
"""
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
Compile Arduino sketch and return hex/binary in a single response.
Synchronous path: held open until the build finishes. Works for AVR /
RP2040 builds (seconds), but ESP-IDF cold builds can run 5-7 minutes
and will hit Cloudflare's 100s edge timeout (HTTP 524). Use the async
path (`/compile/start` + `/compile/status/{job_id}`) for those.
Accepts either `files` (multi-file) or legacy `code` (single file).
Auto-installs the required board core if not present.
"""
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
files = _resolve_files(request)
2026-04-26 05:46:52 +07:00
started = time.monotonic()
try:
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
response = await _run_compile(request, files)
except Exception as e:
2026-04-26 05:46:52 +07:00
await record_compile(
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
user_id=user_id,
2026-04-26 05:46:52 +07:00
project_id=request.project_id,
board_fqbn=request.board_fqbn,
success=False,
duration_ms=int((time.monotonic() - started) * 1000),
error_kind="exception",
extra={"file_count": len(files), "exception": str(e)[:200]},
request=http_request,
)
raise HTTPException(status_code=500, detail=str(e))
2026-04-26 05:46:52 +07:00
duration_ms = int((time.monotonic() - started) * 1000)
await record_compile(
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
user_id=user_id,
2026-04-26 05:46:52 +07:00
project_id=request.project_id,
board_fqbn=request.board_fqbn,
success=response.success,
duration_ms=duration_ms,
error_kind=None if response.success else _classify_compile_error(response.stderr, response.error),
extra={"file_count": len(files), "has_wifi": response.has_wifi},
request=http_request,
)
return response
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
class CompileStartResponse(BaseModel):
job_id: str
class CompileStatusResponse(BaseModel):
state: str # 'pending' | 'running' | 'done' | 'error'
started_at: float
finished_at: float | None = None
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
# Live build output. Grows line-by-line during state=running so the
# frontend can stream it into the compilation console instead of
# waiting for everything to land at the end. Capped at 256 KB
# (most recent tail kept).
stdout: str = ""
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
result: CompileResponse | None = None
error: str | None = None
@router.post("/start", response_model=CompileStartResponse)
async def compile_start(
request: CompileRequest,
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
user_id: str | None = Depends(get_current_user_id),
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
):
"""
Queue a compile and return a `job_id` immediately.
The actual compile runs in a background task; clients then poll
`GET /compile/status/{job_id}` every couple of seconds until state is
`done` or `error`. This sidesteps Cloudflare's 100s HTTP edge timeout —
each individual request returns in milliseconds.
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
Deduplication: identical (files, board_fqbn) submissions while a
matching job is still pending or running return the existing job_id
instead of spawning a new build. Prevents the "user clicks compile six
times six concurrent ninja processes peeling each other apart"
failure mode.
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
"""
files = _resolve_files(request)
_purge_expired_jobs()
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
key = _job_key(files, request.board_fqbn)
existing_id = JOB_BY_KEY.get(key)
if existing_id is not None:
existing = COMPILE_JOBS.get(existing_id)
if existing is not None and existing.get("state") in ("pending", "running"):
logger.info(f"[compile] dedup hit — reusing job {existing_id}")
return CompileStartResponse(job_id=existing_id)
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
job_id = uuid.uuid4().hex
perf(compile): dedup, concurrency limits, and persistent build dir for ESP-IDF Three coordinated fixes that together close the "ESP-IDF compile takes 5-7 min every time" gap and prevent the failure mode where a user clicking compile multiple times spawns six ninja processes that peel each other apart on a modest VPS. What was wrong - /compile/start generated a fresh uuid4 every call, so 6 clicks = 6 independent builds racing each other. Saw load average 30 on the prod VPS during a real BMP280 attempt today. - No concurrency limit anywhere; asyncio.create_task() fired without gating. - ccache was wired in last week (PR #149) but reported 18,350 cacheable calls and **0 hits** because the build dir was a fresh tempfile.TemporaryDirectory(prefix='espidf_') per compile. The random /tmp/espidf_<random>/ path baked into -I and -fmacro-prefix-map flags → different command line every compile → ccache hash miss every time. What this PR does 1. Job deduplication (`backend/app/api/routes/compile.py`) - New `_job_key(files, board_fqbn)` returns SHA-256 of normalised file names + contents + board. Order-independent. - New `JOB_BY_KEY: dict[str, str]` indexes hash → job_id. - `compile_start` checks JOB_BY_KEY before spawning a new task; if a job for this exact content is already pending or running, returns the existing job_id (logs `[compile] dedup hit — reusing job <id>`). - `_purge_expired_jobs` evicts both COMPILE_JOBS and JOB_BY_KEY, keeping the index consistent. Edge case where two jobs share a key (old finished, new running) is handled — only evict the key entry if it still points at the purged job. 2. Concurrency control (`backend/app/api/routes/compile.py`) - `_COMPILE_SEMAPHORE = asyncio.Semaphore(2)` global cap on simultaneous compiles. - `_target_lock(board_fqbn)` returns a per-target asyncio.Lock so concurrent compiles to the SAME board (sharing the persistent build dir) serialise. Different boards still run in parallel up to the semaphore cap. - `_compile_job` acquires sema → per-target lock → flips state to `running` → calls `_run_compile`. Pending state now accurately reflects "queued waiting for resources". 3. Persistent build dir (`backend/app/services/espidf_compiler.py`) - New `_prepare_persistent_project_dir(idf_target)` materialises `/var/lib/velxio-build/<target>/project/` from the template on first use; on subsequent compiles it wipes only `main/` and `user_libs/` (the per-compile parts) and leaves `build/` alone so ninja's incremental cache + ccache .o files survive. - Toolchain version sentinel (`.idf_version`) wipes the whole target dir if the ESP-IDF or arduino-esp32 version changes — cached objects from the old toolchain are no longer ABI-compatible. - `compile()` is now a thin dispatcher: persistent path or fallback to the legacy `tempfile.TemporaryDirectory()` flow. The actual build logic was extracted into `_compile_in_dir()` so both paths share one implementation, no duplication. - Escape hatch: `VELXIO_PERSISTENT_BUILD_DIR=0` env var falls back to the tempfile path without rebuilding the image. Critical for production safety. 4. ccache normalisation (`Dockerfile.standalone`) - + `ENV CCACHE_BASEDIR=/var/lib/velxio-build` makes ccache canonicalise absolute paths under that prefix when computing the cache key. Robustens hits against any future subdir rearrangement. 5. Docker compose (`docker-compose.yml`) - + named volume `velxio-build:/var/lib/velxio-build` so the persistent build dir survives `docker compose up -d --build`. - + env `VELXIO_PERSISTENT_BUILD_DIR=1` (default ON; users disable without rebuilding). Expected impact - Cold first compile per container per target: unchanged (~5-7 min). - Same sketch re-compiled: ~2-5 s (everything cached). - Different sketch, same target: ~5-30 s (only user code + new lib steps rebuild; ESP-IDF base hits cache). - Different sketch with new libraries: ~30-90 s (new lib component compiles; rest hits cache). - Concurrent clicks on same example: 1 build, others poll the same job_id. No more six-ninja meltdown. Tests - `test/backend/unit/test_compile_dedup.py` covers `_job_key` stability + variance and `_purge_expired_jobs` consistency (including the "two jobs share a key" edge case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 13:33:11 +07:00
COMPILE_JOBS[job_id] = {"state": "pending", "started_at": time.time(), "key": key}
JOB_BY_KEY[key] = job_id
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
asyncio.create_task(
_compile_job(
job_id=job_id,
request=request,
files=files,
refactor(oss-split): introduce extension hooks for auth, DB, metrics, auto-save First phase of the OSS / pro split. Goal: open the seams so the auth/DB/admin stack can move into the private overlay (Phase 2-3) without the routes that stay in OSS (compile, libraries, simulation, iot_gateway) having to know. Backend ------- * New app/core/hooks.py — registry for record_compile, get_current_user_id, and lifespan startup tasks. Each hook is a no-op by default; overlays call register_* in register_pro(app) to plug in a real implementation. * compile.py now imports only from app.core.hooks. Drops the direct deps on app.core.dependencies, app.database.session, app.models.user, and app.services.metrics. Route signatures use `Depends(get_current_user_id)` instead of `Depends(get_current_user)`; the metric helper passes user_id through rather than a User instance. * compile_chip.py drops the unused _current_user Depends entirely. * main.py wraps the auth/DB stack import in try/except. When it succeeds (today's behavior on velxio.dev), an adapter bridges record_compile and get_current_user_id to the existing app.services.metrics + dependencies, and the create_all + ALTER TABLE migration block runs via a registered lifespan_startup hook. When it fails (the post-Phase-2 OSS image), main logs "running stateless" and skips registering anything — the routes still load and behave as no-ops for metrics + always-anonymous for auth. Frontend -------- * useAutoSaveProject becomes a skeleton: one useState + one useEffect that delegates to an installed AutoSaveImpl. installAutoSaveImpl() replaces the impl without changing hook count, so React's rules-of-hooks stay satisfied even after the impl moves out of OSS. * New hooks/autoSaveImpl.ts holds the original logic (debouncing, dirty detection, owner eligibility, fetch keepalive on unload), refactored to emit() instead of useState. It self-registers at module load; main.tsx imports it for the side effect. * AppHeader wraps the entire user-vs-login UI in a data-velxio-slot ="header-auth" boundary. Today the OSS UI still renders inside the slot — the overlay can portal-inject additional items now, and in Phase 3 the slot becomes the sole owner of header auth UX. Behavior is identical on velxio.dev (pro overlay imports everything successfully, every adapter wires up). The change is purely structural: deleting the auth/DB modules tomorrow no longer crashes OSS at import. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 23:24:51 +07:00
user_id=user_id,
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
),
)
return CompileStartResponse(job_id=job_id)
@router.get("/status/{job_id}", response_model=CompileStatusResponse)
async def compile_status(job_id: str):
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
"""Poll the status of an async compile job submitted via /compile/start.
`stdout` carries live cmake + ninja output captured line-by-line as
the build runs. Clients should poll every 1-2s and re-render the
full string each time (or compute a length delta). Once state=done,
`result.stdout` carries the same content too both are kept so a
late-arriving poll always has the log available.
"""
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
job = COMPILE_JOBS.get(job_id)
if not job:
raise HTTPException(status_code=404, detail="job not found or expired")
return CompileStatusResponse(
state=job["state"],
started_at=job["started_at"],
finished_at=job.get("finished_at"),
feat(compile): stream live ESP-IDF cmake + ninja output to the console A user reported on Discord: "the Velxio Console doesn't update anything, it just waits until the very end and displays everything in one go". True for the async compile path — /compile/status only carried `state` and the final `result`, so the editor's CompilationConsole stayed empty during the 5-7 minute cold ESP-IDF builds and dumped 1500 lines at once when the build finished. This wires live build output through the whole stack. Backend (espidf_compiler.py) - New _run_with_streaming() helper. When a progress_callback is provided it spawns the subprocess via Popen + stdout/stderr drain threads and invokes the callback line-by-line. When None it falls back to the existing subprocess.run(capture_output=True) one-shot path so the unit-test code that doesn't care about live output is unaffected. - compile() and _compile_in_dir() take an optional ProgressCallback. - _run_cmake / _run_ninja closures now go through _run_with_streaming with that callback. cmake configure (~2-5 s) + ninja (~5-300+ s) both stream now; the ninja output is the one users actually want to watch. Backend (compile.py) - _compile_job seeds COMPILE_JOBS[id]['stdout_buffer'] = '' and defines on_progress_line(line) which appends to it. Buffer capped at 256 KB (tail kept) so a runaway build can't OOM the FastAPI process. - The buffer is preserved on both the success and the error path so late polls still see the log even after state transitions to done/error. - /compile/status now returns the buffer as a `stdout` field. CompileStatusResponse gains the field with default '' so old clients that don't read it still work. Frontend (compilation.ts) - compileCode() takes a 4th argument: optional CompileProgress callback fired every poll while state ∈ {pending, running}. Carries the cumulative stdout (caller computes deltas) plus elapsed seconds. - Surfaces the new `stdout` field of /compile/status and forwards it to the callback. Errors thrown from the callback are swallowed — a faulty UI hook must never break the polling loop. Frontend (EditorToolbar.tsx) - Both compileCode() call sites (Run and Compile-All) now pass an onProgress callback. It tracks `lastStreamedLen` per-compile, splits each new delta on newlines, and appends them as `info`-typed CompilationLog entries via setCompileLogs. The Compile-All flow prefixes each line with the board label so multi-board builds stay readable. - After the build settles, the existing parseCompileResult call still runs and appends the structured analysis on top of the live stream — that's where FAILED-block detection + the `error`-typed entries that drive the auto-switch-to-errors filter live. Net effect on the user complaint: cold ESP-IDF builds now show the ninja [N/1483] progress lines streaming into the console as they happen, instead of staring at an empty panel for 5-7 minutes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 04:36:58 +07:00
stdout=job.get("stdout_buffer", "") or "",
feat(compile): async compile + status polling — no more 524 timeouts The synchronous /api/compile endpoint forced one long-lived HTTP request to span the entire build. Cloudflare's 100s edge timeout cuts that off mid-flight for any cold ESP-IDF compile (BMP280 takes 5-7 min on first run). The user-visible symptom was HTTP 524 well before the backend even noticed. Backend (compile.py) - New `POST /api/compile/start` returns `{job_id}` immediately and spawns the actual compile as an asyncio.create_task background. - New `GET /api/compile/status/{job_id}` returns the current job state (`pending` | `running` | `done` | `error`). Each poll completes in milliseconds, far under any edge timeout. - Existing `POST /api/compile/` kept verbatim for backward compatibility (AVR/RP2040 builds finish in seconds and don't trip 524). - Build logic extracted into `_run_compile()` so both paths share one implementation; no duplicated ESP-IDF / arduino-cli branching. - Async path opens its own short-lived DB session via AsyncSessionLocal for metric recording — the request-scoped session is dead by the time the background task finishes. - COMPILE_JOBS dict purges entries 30 minutes after completion so a busy server doesn't grow unboundedly. Frontend (compilation.ts) - compileCode() now: POST /compile/start → poll /compile/status every 2s until state ∈ {done, error}, with a 15-minute client-side cap. - 30s axios timeout per individual call (not per build) so transient network blips during a long compile auto-retry instead of failing. - 404 on /status throws (job expired / server restarted); other poll errors warn and retry. Surfaces structured error responses verbatim so the editor's compile-error panel keeps working unchanged. Limitation: COMPILE_JOBS lives in-process; if velxio ever scales to multiple FastAPI workers this needs to move to Redis or sqlite. Single- instance is fine today. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:22:09 +07:00
result=job.get("result"),
error=job.get("error"),
)
@router.get("/setup-status")
async def setup_status():
return await arduino_cli.get_setup_status()
@router.post("/ensure-core")
async def ensure_core(request: CompileRequest):
fqbn = request.board_fqbn
result = await arduino_cli.ensure_core_for_board(fqbn)
return result
@router.get("/boards")
async def list_boards():
boards = await arduino_cli.list_boards()
return {"boards": boards}