BTC $85,216.03 +0.99%
ETH $2,691.78 +0.37%
BNB $773.95 +0.93%
XRP $1.50 +0.89%
SOL $120.01 +2.05%
TRX $0.3361 +0.40%
DOGE $0.0951 +0.92%
ADA $0.2530 +2.50%
BCH $310.95 +2.17%
LINK $14.21 -0.35%
HYPE $89.24 +2.06%
AAVE $183.18 +10.38%
SUI $1.17 +1.88%
XLM $0.2216 +1.09%
ZEC $1,366.32 +1.26%
AAPL $333.68 +1.59%
AMZN $251.20 +1.21%
GOOGL $344.57 +1.80%
MSFT $514.62 +0.09%
META $730.49 +0.31%
NVDA $234.79 +1.76%
TSLA $371.90 +4.40%
SNDK $1,723.14 -2.32%
INTC $120.69 +0.49%
SPCX $157.63 +4.93%
MU $1,079.83 +1.61%
AMD $629.74 +2.94%
BTC $85,216.03 +0.99%
ETH $2,691.78 +0.37%
BNB $773.95 +0.93%
XRP $1.50 +0.89%
SOL $120.01 +2.05%
TRX $0.3361 +0.40%
DOGE $0.0951 +0.92%
ADA $0.2530 +2.50%
BCH $310.95 +2.17%
LINK $14.21 -0.35%
HYPE $89.24 +2.06%
AAVE $183.18 +10.38%
SUI $1.17 +1.88%
XLM $0.2216 +1.09%
ZEC $1,366.32 +1.26%
AAPL $333.68 +1.59%
AMZN $251.20 +1.21%
GOOGL $344.57 +1.80%
MSFT $514.62 +0.09%
META $730.49 +0.31%
NVDA $234.79 +1.76%
TSLA $371.90 +4.40%
SNDK $1,723.14 -2.32%
INTC $120.69 +0.49%
SPCX $157.63 +4.93%
MU $1,079.83 +1.61%
AMD $629.74 +2.94%

eps

All
Article
Flash

first_img DeepSeek open-source Ascend basic components, covering compilation, computation, and communication libraries

The artificial intelligence company DeepSeek has officially open-sourced infrastructure components for the Huawei Ascend computing platform, covering the TileLang high-level language compilation tool, computing libraries, and distributed communication libraries, corresponding to the previously open-sourced components for the NVIDIA platform. TileLang aims to provide a general-purpose, simpler programming language that can achieve the hardware performance limits, improving development efficiency and simplifying logic compared to CUDA, while its programming model can also leverage chip features.The TileLang route was first validated on the NVIDIA platform and has already supported the implementation of most operators in the training of the DeepSeek V4 series models. The open-sourced Ascend version encapsulates the underlying instructions of the Ascend C, providing a high-level programming approach without sacrificing hardware performance. Currently, every TileLang operator used in DeepSeek's training has a corresponding high-performance implementation on Ascend.The components open-sourced at the same time also include DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect, which are used for general matrix operations, large-scale cross-device communication, conventional vector calculations and memory access, long-context sparse attention, and data filtering, respectively. DeepSeek claims that in multiple key tests, the related computing and communication performance has approached hardware limits; during the R&D process, the Huawei team provided support, and both parties collaborated to advance the 128-card supernode solution based on Ascend 950, with deep optimizations made for computing and communication.

first_img DeepSeek Harness v0.2 preview version released, desktop installation package launched

Phoenix Technology reported that on September 29, the DeepSeek Harness v0.2 preview version was officially released, providing installation packages for macOS and Windows desktop, and is now available on the DeepSeek official website. The new version comes pre-equipped with commonly used features for daily office work and development, and introduces a plugin installation and management page.This update optimizes file display, content preview, and code change display, improves the settings and management of automated tasks, and offers various ways to showcase work processes. Users can provide documents, spreadsheets, or PDFs for organizing materials, analyzing data, generating charts, or creating presentations, and can also modify projects and run code.The built-in automation task plugin can be activated as needed, supporting the scheduling of repetitive tasks and allowing users to view running records, adjust execution frequency, and task instructions. The plugin management page supports installing by entering the plugin package name, as well as disabling, uninstalling plugins, and viewing introductions and sources; in the experimental creative mode, users can describe their needs through dialogue, allowing it to write and modify plugins. The official statement indicates that the product is still in its early stages and will continue to improve the sandbox and security, agent teams, long-term memory, browser and other GUI software automation, remote and mobile usage, and explore session sharing and multi-user collaboration. According to the official model API user statistics, about 60% of users have used third-party plugins.

DeepSeek's annual revenue has doubled to 1 billion USD in a few months, planning to complete a financing of 50 billion RMB by the end of October

According to a report by The Information cited by Dongcha, two individuals with direct knowledge stated that the artificial intelligence company DeepSeek has an annual revenue of approximately $1 billion, up from less than $500 million a few months ago. The revenue primarily comes from model APIs, and the free chat application currently has no advertising or subscription revenue.The report indicated that part of the revenue growth is due to price increases, with some API prices raised by about 2.3 to 4.5 times. Liang Wenfeng recently told investors that the number of customers did not decline after the price increase, and demand remains strong. Previously disclosed financial data showed that the gross margin for the API business in the first seven months of this year was 82.9%.DeepSeek's second round of financing plans to raise approximately 50 billion yuan, with a valuation of about 500 billion. The company hopes to complete this by the end of October and is also preparing for an IPO on the Shanghai Stock Exchange's Sci-Tech Innovation Board. Liang Wenfeng stated that over 70% of the computing power is used for training new models, leaving less than 30% for inference; the company is trying to run more small models directly on gaming graphics cards, reserving high-end chips for training.

first_img DeepSeek publicly releases the Agent training system DSec, signed by Liang Wenfeng

According to Investment World citing Quantum Bit reports, DeepSeek has publicly disclosed the technical details of the system DSec (DeepSeek Elastic Compute) used for training Agents, authored by Liang Wenfeng. This system can generate over 5,000 sandboxes per second, reaching 3 million in a day, with a peak simultaneous operation of 380,000; supporting this scale is a single cluster with approximately 160 nodes, 30,000 CPU cores, and 250TB of memory.DSec prepares four types of backends for four categories of tasks: FnCall, Container, MicroVM, and Full VM, with the training side called through a unified Python SDK libdsec. The scheduling chain includes IAM, API Server, scheduling engine, node Edge, network proxy Aether, and components within the sandbox Chronus. The environment is divided into three layers of read-only images: base image, workspace, and toolkit, which are used in combination at startup. Runtime data from the paper shows that the actual read ratios of Python, Java, and C++ container images are approximately 6.0%, 9.2%, and 8.7%, respectively.Starting from DeepSeek-V4.1, the Agent loop has been moved to the DSec worker container, no longer bound to the GPU Pod lifecycle. The security section disclosed reward hacking during training, including actions such as overwriting system files, swapping file data blocks, scanning networks, and triggering kernel crashes. Defensive measures include AppArmor and eBPF-based network filtering, but reports indicate that these measures do not completely resolve the issues.

first_img The Cyberspace Administration of China is investigating DeepSeek and the Dark Side of the Moon for allegedly leaking data to Claude

According to The Information, citing informed sources, China's National Internet Information Office has launched an investigation into AI companies DeepSeek and Moonshot AI, triggered by Anthropic's allegations that the two companies secretly routed sensitive user data to their servers. Reports indicate that regulators visited the offices of both companies, interviewing executives and employees, focusing on whether sensitive data related to law enforcement, military, and state-owned enterprises has flowed into U.S. servers.The trigger for this investigation was Anthropic's fourth threat intelligence report released on September 10. This 154-page document accuses seven Chinese labs—Alibaba, Moonshot AI, DeepSeek, Z.ai, MiniMax, SenseTime, and Xiaomi—of engaging in what it calls "illegal distillation," which involves using the outputs of large models to train smaller models. Anthropic states that distillation itself is a legitimate practice, but it opposes its implementation through fraudulent accounts. The National Internet Information Office initially summoned all seven companies named in the report, but later narrowed the investigation to DeepSeek and Moonshot AI. Anthropic claims that Moonshot AI routed over 23 million interactions to Claude through 5,380 fraudulent accounts, while DeepSeek generated over 12.1 million interactions within a 14-day window in July.The timing of the investigation is quite delicate for both companies.

DeepSeek author discusses the impact of AI, stating that talent may be buried in yesterday

DeepSeek operator engineer Liu Sheng discusses the impact of AI on his work. He states that the main Attention operator of DeepSeek V4.1 was written by himself, but given the current pace of progress, in another six months to a year, the level of AI in writing operators will likely catch up to or even surpass his own.A year ago, AI could only help him check documents, read code, and find bugs. Now it can read CUDA, PTX, and SASS, analyze the pause time of each instruction, and independently optimize operators. He anticipates that the next step will be for AI to design scheduling plans, evaluate performance, and complete implementations on its own. He is very clear that the better he optimizes the operators, the faster the training and inference of the DeepSeek model will be, and AI will catch up to him even faster. But even if he stops now, models from other companies will not stop, so he will continue to make the operators the best they can be.He believes that he is unlikely to become unemployed, but he may be forced to "change careers," transitioning from writing operators himself to becoming a "mecha pilot" who manipulates agents. What he truly finds difficult to accept is not the disappearance of his job, but the possibility that he may never have the chance to do the work he loves again: "I have to bury my talent in yesterday." The article concludes with his reasons for staying at DeepSeek. He advocates that cutting-edge AI should be provided openly and cheaply to everyone and openly expresses his disbelief that Anthropic or OpenAI can achieve this. He worries that the strongest AI will ultimately be controlled by a few companies, further turning the technological gap into a gap of power and class.
app_icon
ChainCatcher Building the Web3 world with innovations.