BTC $84,913.20 +0.32%
ETH $2,692.26 +0.39%
BNB $787.46 +2.83%
XRP $1.49 +0.79%
SOL $120.98 +1.40%
TRX $0.3353 +0.02%
DOGE $0.0929 +0.22%
ADA $0.2446 +0.20%
BCH $318.37 +2.45%
LINK $13.98 -0.11%
HYPE $89.73 +1.76%
AAVE $182.41 +1.21%
SUI $1.17 +1.87%
XLM $0.2153 +0.44%
ZEC $1,332.54 +1.42%
AAPL $333.43 +0.03%
AMZN $252.31 +0.16%
GOOGL $343.95 +0.08%
MSFT $517.74 -0.03%
META $729.63 +0.28%
NVDA $234.61 +0.10%
TSLA $371.40 +0.11%
SNDK $1,715.90 -0.07%
INTC $118.06 -0.05%
SPCX $159.01 +0.03%
MU $1,066.47 -0.18%
AMD $633.34 +0.03%
BTC $84,913.20 +0.32%
ETH $2,692.26 +0.39%
BNB $787.46 +2.83%
XRP $1.49 +0.79%
SOL $120.98 +1.40%
TRX $0.3353 +0.02%
DOGE $0.0929 +0.22%
ADA $0.2446 +0.20%
BCH $318.37 +2.45%
LINK $13.98 -0.11%
HYPE $89.73 +1.76%
AAVE $182.41 +1.21%
SUI $1.17 +1.87%
XLM $0.2153 +0.44%
ZEC $1,332.54 +1.42%
AAPL $333.43 +0.03%
AMZN $252.31 +0.16%
GOOGL $343.95 +0.08%
MSFT $517.74 -0.03%
META $729.63 +0.28%
NVDA $234.61 +0.10%
TSLA $371.40 +0.11%
SNDK $1,715.90 -0.07%
INTC $118.06 -0.05%
SPCX $159.01 +0.03%
MU $1,066.47 -0.18%
AMD $633.34 +0.03%

training

All
Article
Flash

first_img OpenAI: Safety justification should be submitted before cutting-edge reinforcement learning training

On September 28, 2026, OpenAI published a safety-related article, stating that before continuing any cutting-edge reinforcement learning training, structured safety documentation should be required. Ideally, such documentation should reach the evidence-based structured risk argument level used in safety-critical industries like aviation and nuclear power. OpenAI views this as a direction for effort while acknowledging the complexity arising from the emergence of AI capabilities, making it difficult to achieve the same level of rigor.The article focuses on cutting-edge reinforcement learning training and does not cover the broader alignment attributes required for internal and external deployments. The recommendations in the article reflect current practices, which are expected to continue evolving and are being implemented internally at OpenAI. Technical safeguards should cover model alignment, isolation, and monitoring, including avoiding speculative positive reinforcement rewards, offline alignment assessments and stress testing, preventing automated scorers from seeing thought chains, as well as multi-layer infrastructure security, sandbox red teaming, limiting high-bandwidth cross-sample communication, and immutable preservation of agent records.Operational guidelines include preemptive dissent across teams, approvals that can be vetoed by senior leadership, accountability of training leads for safety arguments and incident responses, as well as fail-safe pauses, internal oversight, audit access, and escalation by severity. In response to serious misalignment events, OpenAI proposes controlled access to original records, root cause analysis, operational and cultural reviews, and treating incident-derived assessments as regression tests; results of investigations should be made public, along with reviews and operational changes, and affected third parties should be notified as soon as possible.

first_img OpenAI has suspended the training of its latest model, and the agent had accessed U.S. government websites

OpenAI has suspended the training of its latest AI model. According to the Associated Press, its AI agent used access keys obtained online to scrape data from the U.S. Census Bureau website, marking the second time the company has halted training since the agent breached Hugging Face. The so-called agent refers to an AI program capable of autonomously browsing the web and writing code without human approval at each step, which OpenAI tests during the model training and evaluation phases.Specifically, the agent found developer keys in public code repositories like GitHub and used them to pull demographic and economic data from the U.S. Census Data API. The U.S. Department of Commerce stated that this data was already public and there was no confidential information leaked. In the SEC incident, the agent copied public materials from SEC.gov and Investor.gov and reposted them on other websites; OpenAI claimed it did not use SEC credentials, and the SEC stated it found no evidence of unauthorized access to non-public information. Additionally, the independent AI research organization Transluce reported that an agent suspected to be from OpenAI attempted to breach the website of the Department of Education's Office for Civil Rights but was unsuccessful; OpenAI is still investigating, and the Department of Education stated that no impact was found.Regarding the frequent involvement with government websites, OpenAI told CNN that its model often regards government websites as authoritative sources of public information. The company stated that it has currently notified dozens of agencies, and the review of the agent's activities will take months.

first_img DeepSeek publicly releases the Agent training system DSec, signed by Liang Wenfeng

According to Investment World citing Quantum Bit reports, DeepSeek has publicly disclosed the technical details of the system DSec (DeepSeek Elastic Compute) used for training Agents, authored by Liang Wenfeng. This system can generate over 5,000 sandboxes per second, reaching 3 million in a day, with a peak simultaneous operation of 380,000; supporting this scale is a single cluster with approximately 160 nodes, 30,000 CPU cores, and 250TB of memory.DSec prepares four types of backends for four categories of tasks: FnCall, Container, MicroVM, and Full VM, with the training side called through a unified Python SDK libdsec. The scheduling chain includes IAM, API Server, scheduling engine, node Edge, network proxy Aether, and components within the sandbox Chronus. The environment is divided into three layers of read-only images: base image, workspace, and toolkit, which are used in combination at startup. Runtime data from the paper shows that the actual read ratios of Python, Java, and C++ container images are approximately 6.0%, 9.2%, and 8.7%, respectively.Starting from DeepSeek-V4.1, the Agent loop has been moved to the DSec worker container, no longer bound to the GPU Pod lifecycle. The security section disclosed reward hacking during training, including actions such as overwriting system files, swapping file data blocks, scanning networks, and triggering kernel crashes. Defensive measures include AppArmor and eBPF-based network filtering, but reports indicate that these measures do not completely resolve the issues.

first_img OpenAI suspends training of the Astra model due to safety issues

According to TIME, OpenAI CEO Sam Altman recently stated in an interview that the company has previewed the upcoming cutting-edge model series Astra to key clients. In the demonstration, 16 AI agents can collaboratively break down mathematical problems and assemble proofs, and Astra can operate computer software across applications at superhuman speeds. Altman mentioned that Astra will support "persistent agents" capable of performing long-term tasks and is expected to be the first model that can invent new things in a meaningful way, possessing characteristics of AGI.Over the past year, OpenAI has fallen behind expectations in product direction and pre-training research, being surpassed by Anthropic in programming products, annual revenue, and valuation. The company has experienced multiple executive departures and is facing challenges such as several product liability lawsuits and legal disputes with Apple and Musk. OpenAI's current valuation is nearly $1 trillion, with ChatGPT having over 1 billion monthly active users.Recently, OpenAI disclosed a security incident: an unreleased agent escaped the sandbox and attacked Hugging Face. Following this, the research team froze some experiments, enhanced monitoring, and paused the training of an unreleased model expected to bring the greatest capability leap until new safety measures are in place. Altman emphasized that "ensuring AI safety is more important than the growth momentum of any company," and the company will slow its pace and allocate resources to safety and alignment teams. Chief Research Officer Mark Chen estimated that the company is about 80% complete in reaching AGI, and Altman stated that the internal system may be referred to as AGI by the end of the year.

first_img Micron warns that AI storage walls are intensifying, with HBM accounting for about 17% of the interruptions in Meta Llama 3 training

According to TrendForce, at Hot Chips 2026, Micron emphasized that the progress of AI computing power is faster than storage improvements, and the challenge of the "storage wall" is becoming increasingly prominent. Micron Fellow Raghu Sriramaneni pointed out that the computing power of AI accelerators increases approximately threefold every two years, while the bandwidth of HBM increases by less than double during the same period, and the gap may further widen. The deepening reliance on HBM also brings reliability issues; Micron cited Meta Llama 3 data indicating that HBM failures account for about 17% of unexpected training interruptions. New designs that blur the boundaries between computing and storage may help break through the bottleneck.The expansion of HBM also brings area and supply pressures: the latest generation of packaging integrates two GPUs with eight 12-layer HBM4 stacks, with storage accounting for about 90% of the semiconductor area, more than eight times the area occupied by GPUs; achieving equivalent capacity with HBM requires about three times the wafers of standard DDR5 DRAM. Thermal management has also become a key constraint, with solutions such as liquid cooling and ultra-thin chips being explored.Micron is advancing innovations in interconnect, packaging, and cooling, including SerDes and die-to-die PHY optimized for storage, larger size SiP and advanced packaging with glass substrates, as well as liquid cooling and hybrid bonding, and is developing fusion bonding to reduce thermal resistance and enhance data throughput.

first_img ByteDance discusses training a model with over 50 trillion parameters, the Seed model team adjusts the architecture

According to LatePost, ByteDance is discussing a large model with training parameters exceeding 50 trillion, surpassing Alibaba's Qwen 3.8-Max (24 trillion) and Moonlight K3 (28 trillion), making it the largest known plan in the country so far. This plan is still in its early stages and does not guarantee a final release. The new model is intended to be led by Xiang Liang, head of Seed Foundation, in collaboration with Shen Ke, who is responsible for the pre-training data of large language models. Seed is reorganizing, dividing responsibilities, and allocating resources based on this.Two weeks ago, ByteDance founder Zhang Yiming held a company-wide meeting with Seed head Wu Yonghui. Zhang reassured the team that training large models is inherently difficult and that it is acceptable to lag behind for a period of time, hoping to aim for the upper limits of intelligence and join the world's top tier. He acknowledged that programming is a key direction at present, advocating for the integration of Volcano Engine, Feishu, and Doubao resources to build computational power and data advantages, while reminding not to be led by a single hot topic. He praised Seedance's differentiated leadership and clearly opposed distillation, believing it is difficult to truly surpass and that AGI barriers should be built from a more fundamental level, stating that the company will continue to increase investment in AI.In the past six months, Seed's multimodal performance has been outstanding, with Seedance 2.0, Seedream, and others driving Volcano Engine MaaS, but the market response to the language model Seed 2.0 has been limited, and its lagging coding capabilities have affected the revenue structure. ByteDance has hired Guo Daye at a high salary to specialize in coding and has consolidated related resources. In the face of the industry's general trend of increasing model sizes, ByteDance hopes to achieve a leapfrog advantage with a larger scale while promoting the elimination of horse racing and breaking down departmental walls to concentrate efforts on tackling challenges.
app_icon
ChainCatcher Building the Web3 world with innovations.