ilessio-aiflowlab commited on
Commit
29d6413
Β·
verified Β·
1 Parent(s): 95dc957

Upload folder using huggingface_hub

Browse files
Files changed (2) hide show
  1. README.md +163 -0
  2. model.safetensors +3 -0
README.md ADDED
@@ -0,0 +1,163 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: depth-anything/Depth-Anything-V2-Large
4
+ tags:
5
+ - robotics
6
+ - edge-deployment
7
+ - anima
8
+ - forge
9
+ - depth-estimation
10
+ - monocular-depth
11
+ - safetensors
12
+ - vision
13
+ - ros2
14
+ - jetson
15
+ - real-time
16
+ library_name: transformers
17
+ pipeline_tag: depth-estimation
18
+ model-index:
19
+ - name: depth-anything-v2-large
20
+ results:
21
+ - task:
22
+ type: depth-estimation
23
+ metrics:
24
+ - name: Model Size (MB)
25
+ type: model_size
26
+ value: 1279
27
+ ---
28
+
29
+ # Depth Anything V2 Large β€” SafeTensors
30
+
31
+ > Depth Anything V2 (Large, ViT-L backbone) converted to SafeTensors format for safe, fast loading in robotic depth estimation pipelines. 335M parameters for high-quality monocular depth maps.
32
+
33
+ This model is part of the **[RobotFlowLabs](https://huggingface.co/robotflowlabs)** model library, built for the **ANIMA** agentic robotics platform β€” a modular ROS2-native AI system that brings foundation model intelligence to real robots operating in the real world.
34
+
35
+ ## Why This Model Exists
36
+
37
+ Monocular depth estimation is fundamental to robotic navigation and manipulation β€” robots need to know how far away things are from a single camera. Depth Anything V2 produces the highest-quality relative depth maps from a single image. The original weights are distributed as raw `.pth` files. We converted them to SafeTensors format for safe, zero-copy memory-mapped loading.
38
+
39
+ ## Model Details
40
+
41
+ | Property | Value |
42
+ |----------|-------|
43
+ | **Architecture** | DPT head + ViT-Large encoder |
44
+ | **Parameters** | 335M |
45
+ | **Encoder** | ViT-L/14 (DINOv2-based) |
46
+ | **Input Resolution** | Flexible (recommended 518Γ—518) |
47
+ | **Output** | Dense relative depth map |
48
+ | **Training** | Synthetic + real depth labels (multi-stage) |
49
+ | **Original Model** | [`depth-anything/Depth-Anything-V2-Large`](https://huggingface.co/depth-anything/Depth-Anything-V2-Large) |
50
+ | **License** | Apache-2.0 |
51
+
52
+ ## Included Files
53
+
54
+ ```
55
+ depth-anything-v2-large/
56
+ β”œβ”€β”€ model.safetensors # 1.3 GB β€” Full model weights
57
+ └── README.md # This file
58
+ ```
59
+
60
+ ## Quick Start
61
+
62
+ ```python
63
+ from safetensors.torch import load_file
64
+ import torch
65
+
66
+ # Load SafeTensors weights
67
+ state_dict = load_file("model.safetensors")
68
+
69
+ # Load into Depth Anything V2 architecture
70
+ from depth_anything_v2.dpt import DepthAnythingV2
71
+
72
+ model = DepthAnythingV2(encoder='vitl', features=256, out_channels=[256, 512, 1024, 1024])
73
+ model.load_state_dict(state_dict)
74
+ model.to("cuda").eval()
75
+
76
+ # Predict depth
77
+ depth = model.infer_image(image) # Returns relative depth map
78
+ ```
79
+
80
+ ### With Transformers
81
+
82
+ ```python
83
+ from transformers import AutoModelForDepthEstimation, AutoImageProcessor
84
+ import torch
85
+
86
+ processor = AutoImageProcessor.from_pretrained("depth-anything/Depth-Anything-V2-Large")
87
+ model = AutoModelForDepthEstimation.from_pretrained("depth-anything/Depth-Anything-V2-Large")
88
+ model.to("cuda").eval()
89
+
90
+ inputs = processor(images=image, return_tensors="pt").to("cuda")
91
+ with torch.no_grad():
92
+ depth = model(**inputs).predicted_depth
93
+ ```
94
+
95
+ ### With FORGE (ANIMA Integration)
96
+
97
+ ```python
98
+ from forge.vision import VisionEncoderRegistry
99
+
100
+ depth_estimator = VisionEncoderRegistry.load("depth-anything-v2-large")
101
+ depth_map = depth_estimator(image_tensor) # Relative depth map
102
+ ```
103
+
104
+ ## Use Cases in ANIMA
105
+
106
+ Depth estimation is critical across ANIMA modules:
107
+
108
+ - **Obstacle Avoidance** β€” Real-time depth maps for safe navigation
109
+ - **Grasp Planning** β€” Estimate object distance for manipulation reach calculations
110
+ - **3D Reconstruction** β€” Dense depth for point cloud generation from single camera
111
+ - **Safety Zones** β€” Distance-based safety boundaries for human-robot collaboration
112
+ - **Path Planning** β€” Identify traversable spaces and obstacle heights
113
+
114
+ ## Depth Anything V2 Family
115
+
116
+ | Model | Params | Size | Best For |
117
+ |-------|--------|------|----------|
118
+ | **[depth-anything-v2-large](https://huggingface.co/robotflowlabs/depth-anything-v2-large)** | **335M** | **1.3 GB** | **Highest quality depth** |
119
+ | [depth-anything-v2-small](https://huggingface.co/robotflowlabs/depth-anything-v2-small) | 24.8M | 95 MB | Real-time edge deployment |
120
+
121
+ ## Intended Use
122
+
123
+ ### Designed For
124
+ - Monocular depth estimation for robotic navigation
125
+ - Dense depth maps for manipulation planning
126
+ - Point cloud generation from RGB cameras
127
+ - Obstacle detection and distance estimation
128
+
129
+ ### Limitations
130
+ - Produces relative (not metric) depth β€” requires calibration for absolute distances
131
+ - Performance degrades on reflective, transparent, or textureless surfaces
132
+ - Single-frame estimation β€” no temporal consistency for video
133
+ - Inherits biases from training data distribution
134
+
135
+ ### Out of Scope
136
+ - Safety-critical autonomous driving without additional validation
137
+ - Medical depth estimation
138
+ - Surveillance applications
139
+
140
+ ## Attribution
141
+
142
+ - **Original Model**: [`depth-anything/Depth-Anything-V2-Large`](https://huggingface.co/depth-anything/Depth-Anything-V2-Large) by TUM & HKU
143
+ - **License**: [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0)
144
+ - **Paper**: [Depth Anything V2](https://arxiv.org/abs/2406.09414) β€” Yang et al., 2024
145
+ - **Converted by**: [RobotFlowLabs](https://huggingface.co/robotflowlabs) using [FORGE](https://github.com/robotflowlabs/forge)
146
+
147
+ ## Citation
148
+
149
+ ```bibtex
150
+ @article{yang2024depth_anything_v2,
151
+ title={Depth Anything V2},
152
+ author={Yang, Lihe and Kang, Bingyi and Huang, Zilong and Zhao, Zhen and Xu, Xiaogang and Feng, Jiashi and Zhao, Hengshuang},
153
+ journal={arXiv preprint arXiv:2406.09414},
154
+ year={2024}
155
+ }
156
+ ```
157
+
158
+ ---
159
+
160
+ <p align="center">
161
+ <b>Built with FORGE by <a href="https://huggingface.co/robotflowlabs">RobotFlowLabs</a></b><br>
162
+ Optimizing foundation models for real robots.
163
+ </p>
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f045446a0f6f9273f3d3e82fada9ca73d6e01c4229eb88b5c848e94510dce1f5
3
+ size 1341306380