On September 7th, the Solana Foundation introduced the AI project: The project has made approximately 2,000 hours of robot navigation data available, and uses Solana to record data contributions and distribute incentives. It aims to address a real bottleneck in robot training – the high cost and dispersion of high-quality, real-world data, as well as the difficulty of collecting it continuously. In this context, blockchain plays a role in recording contributions, enforcing rules, and distributing value, rather than replacing sensors or automatically verifying the authenticity of each piece of data.
Embodied AI requires the model to understand space, objects, and actions in the real world. Internet text can be replicated on a large scale, but robot data must be collected by devices in specific environments, involving cameras, depth information, pose, and control signals. Different hardware, lighting conditions, and routes can lead to significant variations. BitRobot expands the sources through crowdsourcing and then reduces the cost of repeated data collection for researchers by making the data available openly.
2000 hours to achieve a certain scale, but quality still relies on screening.
"2000 hours" sounds like a lot, but the value of the data cannot be judged solely by duration. Robots may move back and forth in similar corridors repeatedly, or they may record a large number of low-quality segments due to sensor inaccuracies. What truly determines the effectiveness of training are the diversity of scenarios, the consistency of annotations, time synchronization, coverage of failed actions, and whether the hardware metadata is complete.
The advantage of open data is that external teams can re-calculate, clean, and compare models without having to fully rely on the internal results of the project team. Researchers can examine certain approaches, remove damaged frames, and make their own filtering criteria public. However, data openness does not mean there are no licensing restrictions; users still need to check the licenses, privacy requirements, and boundaries for commercial use.
The main benefits of Solana are reflected in its traceable contributor accounts and low-cost payments. When a large number of participants provide small segments of data, traditional platforms need to maintain settlement and regional payment systems; however, on-chain rewards allow contribution records and distribution rules to be integrated into the same process. However, on-chain systems can only prove that a certain address has submitted a particular hash or record, but they cannot verify that the footage has not been forged or that the routes have not been duplicated.
Therefore, the data validation mechanism is more critical than token distribution. Projects need to detect duplicate segments, abnormal sensors, machine-generated fake data, and malicious brushing activities, and link quality ratings to rewards. If payment is based solely on the duration of upload, participants will be incentivized to produce the longest rather than the most useful data.
Privacy is also a practical barrier. When robots navigate in homes, shops, or on streets, they may capture images of faces, license plates, screens, and private spaces. Writing the hashed data onto the blockchain does not mean that the original videos should be made permanently public. Projects need to perform data masking before uploading, establish clear deletion procedures and appeal processes, and avoid directly placing identifiable information into an immutable ledger.
From crowdsourced data to trainable models, there is still a long process involved.
Before data enters the model, it must go through processes such as format unification, calibration, slicing, annotation, and version management. Due to differences in camera perspectives, movement speeds, and control interfaces among different robots, simple concatenation of data may result in the model learning device-specific characteristics rather than general navigation capabilities. Public datasets need to provide sufficient metadata so that users can understand which hardware and environment each piece of data comes from.
Reward design also affects regional distribution. If high-performance robots and stable networks are concentrated in a few areas, the data may over-represent specific buildings, roads, and lifestyles. Projects should track coverage and give greater weight to scarce scenarios, rather than allowing the easiest-to-collect routes to dominate the dataset.
Robot data must also include failures. If only routes that successfully reach the destination are collected, the model will be vulnerable to real-world obstacles, positioning drift, and crowd interference. A high-quality dataset should retain segments such as stops before collisions, re-planning, and manual intervention, and accurately label the reasons for failures. If the reward mechanism punishes all failures, contributors may instead hide the most valuable examples for training.
Hardware calibration determines whether data from different sources can be merged. If the camera's internal parameters, timestamps, odometers, and control frequencies are inconsistent, the same action may be assigned different labels on different devices. BitRobot requires the publication of data specifications and quality tests to allow contributors to identify issues before uploading, as well as to enable downstream researchers to reproduce the processing steps.
The transparency of on-chain mechanisms can facilitate audit rewards, but it also brings issues related to address tracking and speculation. Participants may focus more on token prices rather than studying the value of the tokens, and market fluctuations can also affect their willingness to contribute. Stable and predictable contribution standards are more suitable for long-term data infrastructure than short-term high rewards.
For the AI team, it is necessary to establish an independent baseline before using such data. First, train the model in a small number of scenarios to check if it can generalize to new buildings, different lighting conditions, and various robots; then decide whether to expand the scope of training. If increasing the amount of data does not improve performance in new scenarios, it indicates that there are issues with the filtering or overwriting process.
For the Web3 industry, BitRobot provides a more concrete application than just "being on the chain means being trustworthy": it uses a ledger to coordinate decentralized contributors and automates small incentives. Its success should not be judged solely by the number of transactions on the chain, but rather by how many research teams actually adopt it, whether it can improve robot performance, and whether contributors receive fair rewards.
The project currently showcases 2,000 hours of data and a set of crowdsourcing incentive mechanisms, but this does not mean that the global robotics data market has matured. Data quality, privacy, hardware differences, and anti-cheating measures remain major issues. Solana can make recording and settlement more transparent and efficient, but it cannot solve these problems on its own. Only by combining on-chain incentives with strict offline verification, clear licensing, and genuine model effectiveness can crowdsourced data potentially become a long-term public infrastructure for AI.












