Tsinghua Alumni Explore AI Data Engineering at Haitian Ruisheng: A Deep Dive into Industry Practice

Deep News
09/23

On the afternoon of September 22, 2026, the Tsinghua Alumni Association's AI and Big Data Committee hosted the fifth stop of its "Deep Visit to Listed Companies by Tsinghua Alumni" series at Beijing Haitian Ruisheng Science Technology Ltd. (SH: 688787), a leading A-share player in AI training data. The event drew over 30 alumni specializing in artificial intelligence, big data, and data elements, with committee Executive President Ni Ming, Vice Presidents Huang Xiaobing and Zhu Xuguang, and several deputy secretaries-general in attendance. Li Ke, co-founder and CEO of the company and a Tsinghua alumnus, along with Vice President and Chief Linguist Hao Yufeng, joined the discussions on-site.

During his opening remarks, Ni Ming noted that data supply capability is becoming the decisive variable shaping the ceiling of AI performance. High-quality, scenario-specific, and compliant training data serves as a vital foundation for AI technology deployment. He described the AI data sector as a long-track, high-barrier industry requiring sustained commitment and deep cultivation, where success accrues to those with long-term vision. With 21 years of focus in this field and a strong Tsinghua-rooted culture of long-termism, Haitian Ruisheng stands as a benchmark enterprise worthy of study by fellow alumni. He encouraged attendees to make the most of this face-to-face exchange to broaden industry perspectives and uncover collaboration opportunities.

Li Ke delivered a presentation titled "Data Engineering in the AI Era: Changes, Challenges, and Practice," opening with the cyclical relationship between data supply, AI capability formation, and application feedback to illustrate how data engineering bridges AI capabilities and practical value. Drawing on industry research and reports, he outlined four major shifts: expanding technical boundaries accompanied by growing data volumes; a transition in data demand from sheer volume to structural upgrades; a move from labor-intensive to technology- and knowledge-intensive data supply; and a pivot in data production from scale expansion to quality engineering. Addressing challenges such as scarce professional data, complex multimodal processing, and data credibility, he detailed the company's practices in human-machine collaborative production, data processing platforms, global delivery, and full-lifecycle security management.

Hao Yufeng then explored "How High-Quality Datasets Drive the Value Conversion of Data Elements," narrowing the focus to dataset construction and application. He explained the definition and classification of high-quality datasets, walking through the full process from requirement identification through data planning, integration, cleaning, annotation, and model validation. He highlighted key stages including rule-setting, quality review, and delivery acceptance, as well as methods for evaluating data value through model validation. On the practical side, he shared cases such as transforming cultural tourism videos into multimodal Q&A data, using dialect speech data to support model adaptation, and producing real-world interaction data for embodied intelligence—demonstrating how data from diverse scenarios undergoes processing and quality control to support model training and continuous iteration, thereby realizing tangible value in real-world applications.

Liu Chen, deputy secretary-general of the committee, presented "Intelligently Building an AI-Ready Knowledge Foundation." He pointed out that data governance is undergoing a major transformation in the context of large-model deployment: governance targets are expanding from traditional structured data to vast volumes of unstructured documents, while service recipients are shifting from human-oriented BI systems to AI agents, giving rise to practical challenges like knowledge fragmentation and model hallucination. In response, he proposed an ontology-driven knowledge factory solution: establishing a business ontology semantic layer, using large models to extract entities, business relationships, and rules from enterprise documents, and constructing structured knowledge graphs to address the gaps in traditional vector retrieval approaches. He also introduced a self-evolving ontology mechanism to enable continuous system iteration, rapidly converting knowledge frameworks into deployable business intelligence applications.

In the interactive exchange session, alumni engaged in in-depth discussions with the speakers on topics such as data challenges for embodied intelligence, medical data compliance, and the future of the data annotation industry. Participants agreed that embodied intelligence remains in its early stages, with rapid hardware iteration and unresolved technical roadmaps, meaning that a surge in large-scale data has yet to materialize. For high-value sectors like healthcare, enabling data circulation will require integrated approaches combining anonymization and trusted environments, balancing innovation-driven applications with risk and responsibility considerations.

The visit also included a tour where alumni observed achievements including motion-capture sign language datasets, multilingual speech collection, and multimodal full-lifecycle data annotation platforms. Notably, the national standard sign language dataset had previously supported real-time event broadcasting for hearing-impaired audiences during the Beijing Winter Olympics. This visit gave alumni a more concrete understanding of the full AI data engineering process, high-quality dataset construction, and enterprise-level agent data governance, while also establishing channels for future industry-academia-research collaboration. The Tsinghua Alumni Association's AI and Big Data Committee will continue advancing the listed company deep-visit series, aiming to transform the alumni network into a practical platform linking technology and industry.

Founded in 2005, Beijing Haitian Ruisheng Science Technology Ltd. (SH: 688787) is one of China's earliest professional providers of AI training data and currently the only A-share listed company primarily focused on training data. Over 21 years of development, its data services now cover core domains including intelligent speech (speech recognition and synthesis), computer vision, and natural language processing, serving 22 innovative application scenarios such as autonomous driving, content generation, robotics, smart healthcare, smart education, and smart finance. The company's product and service lines span over 300 major languages and dialects worldwide, with 1,900-plus proprietary AI training data products. Its offerings have earned recognition from domestic and international clients including Alibaba, Tencent, Baidu, iFlytek, Hikvision, ByteDance, Microsoft, Amazon, Samsung, the Chinese Academy of Sciences, and Tsinghua University, with a cumulative customer base of nearly 1,300 entities spanning mainstream technology, internet, social, IoT, autonomous driving, and smart finance enterprises, as well as educational research institutions and select government agencies.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10