Speaker
Description
High-quality, diverse annotations remain a bottleneck for deploying computer vision in
forestry, where scenes are highly variable and targets are defined by operational context rather than appearance alone. A preliminary, fully autonomous dataset creation workflow is presented that links harvester-mounted RGB video to StanForD production records. The pipeline uses visual language models (VLMs) to ground operationally relevant targets in each clip, then applies segmentation with tracking to produce temporally consistent, pixel-accurate instance masks. Because VLM outputs can be non-deterministic, the pipeline includes a consensus-based reliability layer with rejection and fallback logic to reduce grounding errors. The system outputs structured training data aligned with operational metadata and preservs traceability to machine events, enabling reproducible regeneration of labels as models and prompts evolve. Early field-case results are reported for standing trees, a key target in cut-to-length operations.
In an initial clearcutting case study in a pine stand with limited undergrowth, one hour of video processing yielded 52 StanForD-linked cutting events, each represented by a 15-second video clip. The pipeline successfully localized, segmented, and tracked the target tree in 40 of 52 clips (77%). Failures were primarily caused by confusion with high stumps, crane components, stones, and other misdetections. The results indicate that context-aware grounding plus tracked segmentation can generate operationally aligned labels with minimal manual intervention in simple harvesting conditions, while exposing specific failure modes that guide future robustness work.
| Keywords | Machine-vision; Dataset; Machine-learning |
|---|