Memorization vs. Generalization: How Supervised Fine-Tuning SFT and Reinforcement Learning RL Shape Foundation Model Learning
Source: MarkTechPost Modern AI systems rely heavily on post-training techniques like supervised fine-tuning (SFT) and reinforcement learning (RL)...
The Allen Institute for AI (AI2) Releases Tülu 3 405B: Scaling Open-Weight Post-Training with Reinforcement Learning from Verifiable Rewards (RLVR) to Surpass DeepSeek V3 and GPT-4o in Key Benchmarks
Source: MarkTechPost Post-training techniques, such as instruction tuning and reinforcement learning from human feedback, have become essential for...
Meta AI Proposes EvalPlanner: A Preference Optimization Algorithm for Thinking-LLM-as-a-Judge
Source: MarkTechPost The rapid advancement of Large Language Models (LLMs) has significantly improved their ability to generate long-form...
Agentic AI: The Foundations Based on Perception Layer, Knowledge Representation and Memory Systems
Source: MarkTechPost Agentic AI stands at the intersection of autonomy, intelligence, and adaptability, offering solutions that can sense,...
From Deep Knowledge Tracing to DKT2: A Leap Forward in Educational AI
Source: MarkTechPost Knowledge Tracing (KT) plays a crucial role in Intelligent Tutoring Systems (ITS) by modeling students’ knowledge...
Baidu Research Introduces EICopilot: An Intelligent Agent-based Chatbot to Retrieve and Interpret Enterprise Information from Massive Graph Databases
Source: MarkTechPost Knowledge graphs have been used tremendously in the field of enterprise lately, with their applications realized...
Open Thoughts: An Open Source Initiative Advancing AI Reasoning with High-Quality Datasets and Models Like OpenThoughts-114k and OpenThinker-7B
Source: MarkTechPost The critical issue of restricted access to high-quality reasoning datasets has limited open-source AI-driven logical and...
Decoupling Tokenization: How Over-Tokenized Transformers Redefine Vocabulary Scaling in Language Models
Source: MarkTechPost Tokenization plays a fundamental role in the performance and scalability of Large Language Models (LLMs). Despite...
Yandex Develops and Open-Sources Perforator: An Open-Source Tool that can Save Businesses Billions of Dollars a Year on Server Infrastructure
Source: MarkTechPost Yandex, a global tech company, develops and open-sources Perforator, an innovative tool for continuous real-time monitoring...
Quantization Space Utilization Rate (QSUR): A Novel Post-Training Quantization Method Designed to Enhance the Efficiency of Large Language Models (LLMs)
Source: MarkTechPost Post-training quantization (PTQ) focuses on reducing the size and improving the speed of large language models...