[Jun 24, 2026] Free Snowflake Certification GES-C01 Official Cert Guide PDF Download
Snowflake GES-C01 Official Cert Guide PDF
NEW QUESTION # 114
A marketing team is analyzing social media comments using Snowflake and wants to categorize them into predefined campaign sentiments (e.g., 'Positive Campaign Engagement', 'Negative Campaign Feedback', 'Neutral Discussion'). They decide to use the SNOWFLAKE. CORTEX. CLASSIFY TEXT function for this task. Which of the following statements about its usage are correct?
- A. The argument must contain exactly two string values for effective binary classification, otherwise an error is returned.
- B. If the input text exceeds a model-specific token limit, CLASSIFY_TEXT will automatically truncate the text before processing without raising an error.
- C. To provide more context and potentially improve classification accuracy, categories within the can be defined as SQL objects, including 'description' and 'examples' fields.
- D. The input string to CLASSIFY_TEXT is case-insensitive, meaning 'Great product!' and 'great product!' will yield identical classification results due to automatic normalization.
- E. CLASSIFY_TEXT can return a JSON object with a 'label' field, where the value of this field indicates the classified category of the input text.
Answer: C,E
Explanation:
Option A is incorrect because both the input string and categories for 'CLASSIFY TEXT are case sensitive, meaning different capitalizations can lead to different results. Option B is incorrect because the argument must contain at least two and at most 100 unique categories. Option C is correct as returns an OBJECT (VARIANT) whose 'label' field specifies the category to which the input prompt belongs. Option D is correct because categories can be simple strings or SQL objects, allowing for a description and examples to be provided, which can improve accuracy. Option E is incorrect because the documentation for 'CLASSIFY _ TEXT does not mention automatic truncation of input text based on a token limit, although LLMs typically have context windows. The source only mentions that for non- plain English text, results may not be what you expect, not that the input would be truncated.
NEW QUESTION # 115
A team is developing a critical business intelligence application that leverages Snowflake Cortex Analyst to provide natural language querying capabilities over complex structured dat a. To minimize operational costs while maintaining high accuracy, which of the following strategies are most effective for optimizing the cost efficiency of the Cortex Analyst service?
- A. Using a smaller, less capable LLM as the underlying summarization agent for multi-turn conversations to reduce token processing costs, even if it slightly degrades conversational context.
- B. Configuring a custom instruction with a short, precise task description to reduce the input token count for the LLMs orchestrating SQL generation.
- C. Implementing a comprehensive Verified Query Repository (VQR) to guide Cortex Analyst towards pre-validated SQL queries for common questions, which ensures predictable execution and reduces LLM inference iterations.
- D. Optimizing the semantic model YAML file by reducing the number of logical tables and columns to decrease the metadata processed by Cortex Analyst's LLMs per message.
- E. Leveraging Cortex Search Services integration within the semantic model to improve literal value matching, thereby reducing the need for Cortex Analyst to perform expensive fuzzy string matching or re-prompt the user.
Answer: C,E
Explanation:
Option B is correct because a Verified Query Repository (VQR) helps Cortex Analyst leverage pre-validated SQL for similar questions, improving accuracy and potentially reducing the number of LLM inference calls or complex reasoning steps required for SQL generation, thus making usage more efficient and reducing cost associated with less optimal LLM calls. Option D is correct because integrating Cortex Search Services improves literal search, helping Cortex Analyst find exact literal values needed for SQL queries more accurately and efficiently, which can reduce ambiguity and the need for multiple LLM iterations or incorrect queries, ultimately leading to more cost-effective message processing. Option A is incorrect: While using a smaller LLM might seem to save cost, Llama 3.1 70B was specifically chosen as the summarization agent for multi-turn conversations in Cortex Analyst due to its higher accuracy in rephrasing questions and avoiding errors, implying that a less capable model would degrade performance and potentially lead to more (and thus more expensive) overall messages to achieve a correct answer. The cost for Cortex Analyst is per message, not per token for this component. Option C is incorrect. While a well- scoped semantic model is recommended for accuracy, the sources do not explicitly state that reducing the number of logical tables and columns 'directly' reduces the per-message cost of Cortex Analyst, which is fixed per message. The impact would be indirect through improved accuracy or reduced processing complexity, but not a direct cost reduction based on metadata size for the fixed per-message billing. Option E is incorrect. Cortex Analyst cost is based on the number of messages, not the token count of prompts. While good prompt engineering (like concise custom instructions) is generally good practice, it does not directly reduce the per-message cost of Cortex Analyst as it would for token- based LLM calls.
NEW QUESTION # 116
A data scientist is designing a real-time similarity search feature in Snowflake using product embeddings. They plan to use VECTOR_L2_DISTANCE to find similar products. Which statement correctly identifies a cost or data type characteristic relevant to this implementation?
- A. The maximum dimension supported for a
- B. The
- C. The
- D. Both the
- E. Storing product embeddings generated by
Answer: C
Explanation:
Option A is incorrect because vector similarity functions, including

NEW QUESTION # 117
A business user frequently asks Cortex Analyst questions that require filtering on specific product names, such as "What were the sales for 'iced tea' last month?" The 'product' dimension has many distinct values (high cardinality), and Cortex Analyst sometimes struggles to accurately identify the exact literal product name, leading to less precise SQL queries. The Gen AI Specialist wants to enhance Cortex Analyst's ability to find these literal values for the 'product' dimension. To improve Cortex Analyst's literal search capability for the high-cardinality 'product' dimension, which of the following is the most appropriate and recommended approach to configure in the semantic model?
- A. Option D
- B. Option A
- C. Option C
- D. Option B
- E. Option E
Answer: D
Explanation:
Cortex Analyst offers solutions to improve literal usage, including semantic search over sample values in the semantic model and semantic search using Cortex Search Services. For dimensions with high cardinality (many distinct values), creating a Cortex Search Service on the underlying column and specifying it in the field of the dimension within the semantic model is the recommended approach. This allows for high-quality "fuzzy" search to find literal values needed for Cortex Analyst's SQL queries. Option A is less effective for high-cardinality dimensions because only a fixed-size set of sample values is presented to the LLM, regardless of how many are provided. Option C is not the intended use for the 'description' field and could exceed context window limits. Option D, while a possible technical solution, bypasses the integrated and optimized Cortex Search functionality designed for this purpose. Option E is explicitly contradicted by the scenario, which indicates the LLM struggles, and the available solutions are designed to address this limitation.
NEW QUESTION # 118
A data scientist is tasked with improving the accuracy of an LLM-powered chatbot that answers user questions based on internal company documents stored in Snowflake. They decide to implement a Retrieval Augmented Generation (RAG) architecture using Snowflake Cortex Search. Which of the following statements correctly describe the features and considerations when leveraging Snowflake Cortex Search for this RAG application?
- A. To create a Cortex Search Service, one must explicitly specify an embedding model and manually manage its underlying infrastructure, similar to deploying a custom model via Snowpark Container Services.
- B. The

- C. Cortex Search automatically handles text chunking and embedding generation for the source data, eliminating the need for manual ETL processes for these steps.
- D. Enabling change tracking on the source table for the Cortex Search Service is optional; the service will still refresh automatically even if change tracking is disabled.
- E. For optimal search results with Cortex Search, source text should be pre-split into chunks of no more than 512 tokens, even when using models with larger context windows like

Answer: B,C,E
Explanation:
Option A is correct because Cortex Search is a fully managed service that gets users started with a hybrid (vector and keyword) search engine on text data in minutes, without needing to worry about embedding, infrastructure maintenance, or index refreshes. Option B is incorrect because Cortex Search is a fully managed service; users do not need to manually manage the embedding model infrastructure. A default embedding model is used if not specified. Option C is correct because, for best search results with Cortex Search, Snowflake recommends splitting text into chunks of no more than 512 tokens, as smaller chunks typically lead to higher retrieval and downstream LLM response quality, even with models that have larger context windows. Option D is correct because the SNOWFLAKE.CORTEX.SEARCH_PREVIEW' function allows users to test the search service to confirm it is populated with data and serving reasonable results for a given query. Option E is incorrect because change tracking is required on the source table for the Cortex Search Service to function correctly and reflect updates to the base data.
NEW QUESTION # 119
A Gen AI Specialist is setting up Snowpark Container Services (SPCS) to host a custom open-source LLM. They need to understand the fundamental nature and constraints of image repositories within Snowflake. Which of the following statements accurately describe image repositories in Snowflake's Snowpark Container Services?
- A. Image repositories are primarily used for storing pre-trained Snowflake ML models and do not support custom third-party container images.
- B. Dropping individual images from a repository is supported, allowing for granular management of stored container images.
- C. An image repository in Snowflake is an OCIv2 compliant service used exclusively for storing Docker images, not other OCI-compliant container images.
- D. Image repositories are storage units within an image registry service, and they are used to store OCI-compliant container images.
- E. The maximum compressed layer size for an image registry is consistent across all cloud providers and regions, typically 160 GiB.
Answer: D
Explanation:
Option D is correct because an image registry is an OCIv2 compliant service for storing OCI-compliant container images, and an image repository is defined as a storage unit within that image registry service. Option A is incorrect as image registries are for all OCI-compliant container images, not exclusively Docker images. Option B is incorrect because dropping individual images from a repository is currently not supported; instead, dropping a repository removes all images within it. Option C is incorrect as the maximum layer size for an image registry varies by cloud provider (160 GiB for AWS and 195 GiB for Azure). Option E is incorrect because Snowpark Container Services is designed to run various containerized workloads, including custom third-party container images and open-source LLMs, providing flexibility beyond Snowflake's pre-trained ML models.
NEW QUESTION # 120
A financial analytics team is developing an application to extract specific, structured financial data (e.g., company name, revenue, profit margin) from various news articles using Snowflake Cortex LLM functions. They require the output to strictly conform to a predefined JSON schema and want to ensure robust error handling. Which of the following statements are crucial considerations for achieving this goal?
- A. Setting the temperature option to 0 in the AI_COMPLETE call is essential for obtaining the most consistent and accurate structured JSON outputs, regardless of task complexity or model used.
- B. To guarantee that critical fields like 'company name' and 'revenue' are always extracted, these properties must be explicitly listed within the ' required' array of the JSON schema provided to AI_COMPLETE.
- C. For enhanced reliability in production pipelines, the team should wrap their AI_COMPLETE calls within TRY_COMPLETE, as it returns a structured error object if the model fails to adhere to the schema, allowing for detailed debugging.
- D. The complexity of the JSON schema provided to AI_COMPLETE has no impact on compute costs, as only the input text and generated content tokens are billed.
- E. The AI_COMPLETE function should be used with the response_format argument, supplying a JSON schema object that defines the required structure, data types, and constraints for the output.
Answer: A,B,E
Explanation:
Option A is correct. AI_COMPLETE Structured Outputs allows specifying a JSON schema via the argument to ensure response _ format responses follow a defined structure, data types, and constraints. Option B is correct. Using the field in the JSON schema ensures that required specified properties are extracted, or an error is raised by making extraction of critical information reliable. Option C is incorrect. COMPLETE, performs the same operation as COMPLETE (or AI_COMPLETE) but returns instead of raising an error when the operation cannot be TRY COMPLETE NULL performed. It does not return a structured error object for detailed debugging, but rather handles the error by returning allowing a pipeline to NULL, continue. Option D is correct. For the most consistent results and to optimize JSON adherence accuracy, it is recommended to set the temperature option to 0 when calling COMPLETE (or AI_COMPLETE). Option E is incorrect. The number of tokens processed (and billed) increases with schema complexity. A larger and more complex supplied schema generally consumes more input and output tokens, leading to higher compute costs.
NEW QUESTION # 121
A Gen AI specialist is preparing to upload a large volume of diverse documents to an internal stage for Document AI processing. The objective is to extract detailed information, including lists of items and potentially classifying document types, and then automate this process. Which of the following statements represent 'best practices or important considerations/limitations' when preparing documents and setting up the Document AI workflow in Snowflake? (Select ALL that apply.)
- A. For continuous processing of new documents, it is best practice to create a stream on the internal stage and a task to automate the '!PREDICT method execution.
- B. When defining data values for extraction, especially for nonstandard formats or combinations of values, fine-tuning the model with annotations is generally more effective than relying solely on complex prompt engineering.
- C. Documents with a page count exceeding 125 pages or a file size greater than 50 MB will be processed, but with a potential reduction in extraction accuracy.
- D. If the Document AI model does not find an answer for a specific field, the '!PREDICT method will omit the 'value' key but will still return a 'score' key to indicate confidence that the answer is not present.
- E. To improve model training, documents uploaded should represent a real use case, and the dataset should consist of diverse documents in terms of both layout and data.
Answer: A,B,D,E
Explanation:
NEW QUESTION # 122
An operations team at a company is implementing a robust governance framework to monitor and optimize the costs associated with their Snowflake Cortex LLM function usage. They need to identify which functions are driving the highest token consumption and overall credit usage to pinpoint areas for cost reduction. Which of the following monitoring tools or methods are appropriate for gaining these insights into Cortex LLM function costs and token consumption?
- A. Option C
- B. Option B
- C. Option A
- D. Option D
- E. Option E
Answer: B,C,D,E
Explanation:
NEW QUESTION # 123
A data scientist fine-tuned a mistral-lb model in Snowflake for a specific customer support response generation task, naming it my_custom_responder_model. They now want to make this model available for AI_COMPLETE calls in production, ensuring proper access control and regional availability. Which of the following statements is true regarding the deployment and management of this fine-tuned model in Snowflake?
- A. Option D
- B. Option A
- C. Option C
- D. Option E
- E. Option B
Answer: C
Explanation:
NEW QUESTION # 124 
- A. The parameter provides granular control at the database or schema level, allowing administrators to define different sets of approved models for different data contexts.
- B. If a user attempts to call an LLM not explicitly listed in the via the Cortex LLM Playground, Snowflake will automatically fall back to the 'snowflake-arctic' model for that request.
- C. Enabling Cortex Guard for 'AI_COMPLETE through the 'guardrails' option automatically bypasses the 'CORTEX_MODELS_ALLOWLIST to ensure all LLMs can be used for content safety filtering.
- D. Users with 'ds_team_role' will still be able to successfully call

- E. The 'ACCOUNTADMIN' must execute the SQL command:

Answer: D,E
Explanation:
NEW QUESTION # 125
A developer has successfully created a Cortex Search Service named transcript _ search service within cortex_search_db. services based on customer support transcripts. They now need to query this service to find support tickets related to 'internet issues' specifically from the 'North America' region, and they only want the single most relevant result. Which of the following SQL commands correctly performs this query?
- A. Option D
- B. Option C
- C. Option A
- D. Option E
- E. Option B
Answer: C
Explanation:
Option A is the correct syntax for querying a Cortex Search Service using the SEARCH PREVIEW function, as demonstrated in the sources. It correctly specifies the service name, query, columns to retrieve, filter condition using the operator, and the limit, all within a JSON @eq string. Options B, C, D, and E use incorrect function names, parameter formats, or JSON structures for the SEARCH PREVIEW function as defined in the provided documentation.
NEW QUESTION # 126
A Gen AI developer has a Document AI pipeline that uses a query with 'GET PRESIGNED URL' to process multi-page PDF documents. Despite the internal stage being correctly set up with 'SNOWFLAKE SSE' encryption and the model build being published, they observe inconsistent results. Some documents result in a Received HTTP 403 response for presigned URL. URL may be expired.
error, while other documents (containing complex diagrams and dense text in an unsupported language like Korean) are processed, but the extracted information is often incomplete or inaccurate.
Which two factors are most likely contributing to these observed issues?
- A. The '!PREDICT method is being called with an outdated model build version instead of the latest one, leading to performance degradation.
- B. The Document AI model is returning answers longer than its limit of 512 tokens for entity extraction or 2048 tokens for table extraction.
- C. The default expiration time for the 'GET PRESIGNED URL' function is causing some URLs to expire before the Document AI model can process them.
- D. The role lacks the 'EXECUTE TASK' privilege, preventing the scheduled pipeline tasks from running.
- E. The documents are in an unsupported language or exceed the maximum page length of 125 pages, causing extraction failures or inaccuracies.
Answer: C,E
Explanation:
The error 'Received HTTP 403 response for presigned URL. URL may be expired.' directly indicates that the function's default expiration time is causing some documents to be inaccessible by the Document AI model. This is a common issue when processing pipelines encounter delays. Additionally, the observation of incomplete or inaccurate extraction for documents with 'dense text in an unsupported language like Korean' directly points to language limitations. Document AI explicitly lists supported languages (English, Spanish, French, German, Portuguese, Italian, and Polish) and states that results for other languages might not be satisfactory. While the question mentions 'multi-page PDF documents' without explicitly stating they exceed page limits, the mention of 'complex diagrams and dense text' can also imply potential issues if page length (max 125 pages) is exceeded or other document requirements are not met. Thus, option D comprehensively covers these content- related issues. Option A (outdated model version) is unlikely to cause these specific errors, as the latest model is used by default if not specified. Option C (missing 'EXECUTE TASK privilege) would prevent task execution entirely, not cause intermittent URL issues or content- specific extraction problems. Option E (answers exceeding token limits) would be reflected in truncated output, not necessarily 'incomplete or inaccurate' extraction in the sense of failing to identify information in the first place.
NEW QUESTION # 127
A data engineering team needs to establish an automated pipeline in Snowflake to continuously extract 'contract_id' and effective_date' from new PDF contract documents uploaded to an internal stage named They have a pre-trained Document AI model named 'contract_processor'. Which of the following sets of SQL commands correctly configures the necessary Snowflake objects for this automated processing pipeline, including handling file access and initial data loading?
- A.

- B.

- C.

- D.

- E.

Answer: A
Explanation:
NEW QUESTION # 128
A financial institution uses Snowflake Cortex LLM functions to process customer feedback. They initially used SNOWF LAKE .CORTEX.SENTIMENT for general sentiment analysis. Now, they need to extract specific sentiment categories (e.g., 'service_quality', 'product_pricing') and the sentiment for each, expecting the output in a structured JSON format for automated downstream processing. Which AI_COMPLETE configuration best addresses their new requirement while considering cost-efficiency and output reliability?
- A.

- B.

- C.

- D.

- E.

Answer: D
Explanation:
Option B is correct. For medium-complexity tasks like extracting specific sentiment categories into a structured format, Snowflake recommends using more powerful models, explicitly prompting the model to 'Respond in JSON', providing detailed descriptions for schema fields, and setting fields as 'required' to improve accuracy and ensure adherence to the schema.
is a smaller model which might struggle with the accuracy and reliability required for complex structured extraction compared to more powerful models, even with a schema. Option C is incorrect because a temperature of 1.0 increases randomness, which is detrimental to the reliability and consistency required for structured JSON output and automated processing. The response_format should also be specified in the options argument explicitly for structured output. Option D is incorrect; while mistral-large2 is a powerful model, relying on guardrails alone does not guarantee structured output or adherence to a specific JSON schema for complex extraction. For complex tasks, explicit prompting and schema details are crucial. Option E is incorrect because
) returns a single classification label and cannot produce a JSON object with multiple specific sentiment categories and their respective sentiments from a single input text as required by the scenario.
NEW QUESTION # 129
A development team is preparing to deploy a new Retrieval-Augmented Generation (RAG) application written in Python. They intend to use Snowflake AI Observability to capture detailed logs and traces for debugging and performance analysis. Which of the following configurations are essential prerequisites for enabling this logging capability effectively?
- A. Option D
- B. Option C
- C. Option B
- D. Option A
- E. Option E
Answer: B,C,D,E
Explanation:
NEW QUESTION # 130
A data architect is integrating Snowflake Cortex LLM functions into various data enrichment pipelines. To ensure optimal performance, cost-efficiency, and accuracy, which of the following are valid best practices or considerations for these pipelines?
- A. When extracting specific entities from documents using SAI EXTRACT or '!PREDICT , it is often more effective to fine-tune a Document AI model for complex or varied document layouts rather than relying solely on extensive prompt engineering for zero-shot extraction.
- B. When performing sentiment analysis on customer feedback using 'AI_SENTIMENT, it's best practice to pass detailed, multi-turn conversation history to the function to enhance accuracy, similar to how 'AI_COMPLETE handles conversational context.
- C. For tasks requiring deterministic JSON outputs, explicitly specifying a JSON schema using the 'response_format' argument with 'AI COMPLETE is crucial, and for OpenAI (GPT) models, including the 'required' field and setting 'additionalPropertieS to 'false' in every node of the schema is a mandatory requirement.
- D. For data enrichment involving classification with 'AI_CLASSIFY' , using descriptive and mutually exclusive categories in plain English, along with an optional clear task description, can significantly improve classification accuracy.
- E. To manage costs effectively for LLM functions like SAI COMPLETE in a pipeline, always use the largest available warehouse size (e.g., 6XL Snowpark- optimized) to maximize throughput, as this directly reduces the overall token processing time and cost.
Answer: A,C,D
Explanation:
Option A is correct. For extracting information from documents with complex or varied layouts, fine-tuning a Document AI model can significantly improve results compared to relying solely on zero-shot extraction and extensive prompt engineering. Document AI provides both zero-shot extraction and fine-tuning capabilities, with fine-tuning recommended to improve results on specific document types. Option B is correct. To ensure 'AI_COMPLETE (or 'COMPLETE) returns responses in a structured JSON format, it is essential to specify a JSON schema using the 'response_format' argument. For OpenAl (GPT) models, specific requirements include setting 'additionalPropertieS to 'false' in every node and ensuring the 'required' field lists all property names. Option C is incorrect. Snowflake explicitly recommends executing queries that call Cortex AISQL functions (such as 'AI COMPLETES) using a smaller warehouse, no larger than MEDIUM. Using larger warehouses does not increase performance for these functions but will incur unnecessary compute costs. The LLM inference itself is managed by Snowflake, and its performance isn't directly scaled by warehouse size in the same way as traditional SQL queries. Option D is incorrect. 'AI_SENTIMENT (and 'SENTIMENT) is a task-specific function designed to return a sentiment score for a given English-language text. Unlike 'AI_COMPLETE (or 'COMPLETE'), which supports multi-turn conversations by passing conversation history for a stateful experience, SAI SENTIMENT processes individual text inputs and is not designed to leverage multi-turn context in the same way for sentiment analysis. Option E is correct. For classification tasks using 'AI_CLASSIFY (or 'CLASSIFY TEXT), best practices include using plain English for the input text and categories, ensuring categories are descriptive and mutually exclusive, and adding a clear 'task_description' when the relationship between input and categories is ambiguous. These guidelines significantly improve classification accuracy.
NEW QUESTION # 131
A data engineering team is setting up an automated pipeline in Snowflake to process call center transcripts. These transcripts, once loaded into a raw table, need to be enriched by extracting specific entities like the customer's name, the primary issue reported, and the proposed resolution. The extracted data must be stored in a structured JSON format in a processed table. The pipeline leverages a SQL task that processes new records from a stream. Which of the following SQL snippets and approaches, utilizing Snowflake Cortex LLM functions, would most effectively extract this information and guarantee a structured JSON output for each transcript?
- A. Option D
- B. Option A
- C. Option C
- D. Option E
- E. Option B
Answer: C
Explanation:
To guarantee a structured JSON output for entity extraction, (the updated version of 'COMPLETE()') with the response_format' argument and a specified JSON schema is the most effective approach. This mechanism enforces that the LLM's output strictly conforms to the predefined structure, including data types and required fields, significantly reducing the need for post-processing and improving data quality within the pipeline. Option A requires multiple calls and manual JSON assembly, which is less efficient. Option B relies on the LLM's 'natural ability' to generate JSON, which might not be consistently structured without explicit 'response_format' . Option D uses , which is for generating summaries, not structured entity extraction. Option E involves external LLM API calls and Python UDFs, which, while possible, is less direct than using native 'AI_COMPLETE structured outputs within a SQL pipeline in Snowflake Cortex for this specific goal.
NEW QUESTION # 132
A data scientist is leveraging various Snowflake Cortex LLM functions to process extensive text data for an application. To effectively manage their budget, they need a clear understanding of how costs are incurred for each specific function. Which of the following statements accurately describe how costs are calculated for Snowflake Cortex LLM functions, with a particular focus on token usage?
- A. Option A
- B. Option C
- C. Option B
- D. Option E
- E. Option D
Answer: C,E
Explanation:
Option B is correct because for the 'EXTRACT ANSWER function, the number of billable tokens is the sum of the tokens in the 'from_text' (source_document) and 'question' fields. Option D is correct as for 'CLASSIFY_TEXT (or 'AI_CLASSIFY), labels, descriptions, and examples provided in the categories are counted as input tokens for each record processed, which directly increases the cost. Option A is incorrect because and functions only count 'input tokens' towards the billable total, not both input and output tokens. Option C is incorrect because Cortex COMPLETE Structured Outputs does not incur additional compute cost for the overhead of verifying tokens against the supplied JSON schema, although schema complexity can increase total token consumption. Option E is incorrect because 'AI_PARSE_DOCUMENT (and 'SNOWFLAKE.CORTEX.PARSE_DOCUMENT) billing is based on the 'number of document pages processed' (e.g., 3.33 Credits per 1,000 pages for Layout mode), not just the number of documents.
NEW QUESTION # 133
A data platform administrator needs to retrieve a consolidated overview of credit consumption for all Snowflake Cortex AI functions (e.g., LLM functions, Document AI, Cortex Search) across their entire account for the past week. They are interested in the aggregated daily credit usage rather than specific token counts per query. Which Snowflake account usage views should the administrator primarily leverage to gather this information?
- A. Option D
- B. Option A
- C. Option C
- D. Option B
- E. Option E
Answer: D
Explanation:
NEW QUESTION # 134
A global analytics firm is developing a Retrieval Augmented Generation (RAG) system in Snowflake to answer customer queries across a large repository of technical documentation, which includes documents in English, German, and Spanish. They are looking to use a Snowflake Cortex embedding model to convert document chunks into vector embeddings for their Cortex Search Service. Which of the following considerations are critical when selecting an appropriate embedding model to optimize for both query relevance and cost-efficiency for their multilingual RAG application? (Select all that apply)
- A. Option D
- B. Option A
- C. Option C
- D. Option B
- E. Option E
Answer: C,D
Explanation:
Option B is correct because both
model provides an increased context window of 8000 tokens while maintaining the same cost per million tokens (0.05 credits) as the 512-token version of
Option C is correct because Snowflake recommends splitting text into chunks of no more than 512 tokens for best search results with Cortex Search, as research shows this typically leads to higher retrieval precision and improved downstream LLM response quality, even when using longer-context embedding models. Option A is incorrect because
is an English-only embedding model, which does not meet the requirement for multilingual documentation. Option D is incorrect; the cost per million tokens for EMBED_TEXT_1024 models (e.g., 0.05-0.07 credits) is not inherently more cost-efficient than EMBED_TEXT_768 models (e.g., 0.03 credits), and cost-efficiency depends on the specific model and use case, not just output dimensions. Option E is incorrect; the context window of an embedding model refers to the maximum length of a text input (chunk) it can process. The maximum pages a document can have (e.g., 300 pages for Document AI) is a separate document requirement, not directly determined by the embedding model's context window.
NEW QUESTION # 135
A data engineering team is planning to build a real-time data pipeline using Snowflake's dynamic tables to process incoming log dat a. They want to use SNOWFLAKE. CORTEX. EXTRACT_ANSWER to pull out specific error codes and timestamps from log entries. They are also mindful of the operational costs. Which of the following statements accurately describes limitations or cost considerations for using SNOWFLAKE . CORTEX. EXTRACT_ANSWER in this scenario?
- A. To optimize performance and reduce cost for large-scale log processing, it is recommended to execute EXTRACT_ANSWER queries on a Snowpark-optimized warehouse of at least 'LARGE' size.
- B. The billing for EXTRACT_ANSWER is based on the combined token count of both the source_document (log entry) and the question used for extraction.
- C. The EXTRACT_ANSWER function does not support dynamic tables, which will prevent its direct use in this pipeline design.
- D. If a source_document input exceeds the model's token limit, EXTRACT_ANSWER will automatically truncate the text, potentially impacting results but avoiding an error.
- E. EXTRACT_ANSWER is explicitly designed for multi-language extraction, making it suitable for logs that might contain various languages without affecting cost.
Answer: B,C
Explanation:
Option A is correct because Snowflake Cortex functions, including , do not support dynamic tables. Option B is correct because for EXTRACT_ANSWER, the number of billable tokens is the sum of the number of tokens in the (source document) and 'question' fields. Option C is incorrect; Snowflake recommends executing queries that call Cortex AISQL functions with a smaller warehouse (no larger than MEDIUM) as larger warehouses do not increase performance. Option D is incorrect; inputs that exceed the model's token limit result in an error, rather than automatic truncation. Option E is incorrect; while the newer 'AI_EXTRACT supports multiple languages, the documentation for "EXTRACT_ANSWER does not explicitly state multi-language support and generally refers to plain English text. Costs are incurred per token regardless of language.
NEW QUESTION # 136
An 'ACCOUNTADMIN' has configured the 'CORTEX MODELS ALLOWLIST parameter to allow only the 'mistral-large? model. A developer, whose role has been granted 'SNOWFLAKE.CORTEX USER and the specific application role 'SNOWFLAKE."CORTEX- MODEL-ROLE-LLAMA3.1-70B"' , subsequently accesses the Cortex LLM Playground. Which models would be available for selection and successful inference by this user within the Playground?
- A. Option D
- B. Option C
- C. Option B
- D. Option A
- E. Option E
Answer: C,D
Explanation:
The parameter governs which models can be used with 'COMPLETE, 'TRY_COMPLETE , and are therefore available in the Cortex LLM Playground. Option A is correct because 'mistral-large? is explicitly in the , making it available. Option B is also correct because Snowflake Cortex functions, including 'AI_COMPLETE (which the Playground utilizes), first check if the provided model name is an identifier for a schema-level model object. If found, role-based access control (RBAC) is applied. Since the user's role has 'SNOWFLAKE."CORTEX-MODEL-ROLE-LLAMA3.1-70B"' granted, they have access to the 'LLAMA3.1-70B" model object, overriding the account-level allowlist for this specific model object. Options C and D are incorrect because snowflake-arctic' is neither in the allowlist nor is specific RBAC granted for it in this scenario. Option E is incorrect as RBAC on model objects can grant access to models not explicitly listed in the account-level when referenced as model objects.
NEW QUESTION # 137
An enterprise is deploying a new RAG application using Snowflake Cortex Search on a large dataset of customer support tickets. The operations team is concerned about managing compute costs and ensuring efficient index refreshes for the Cortex Search Service, which needs to be updated hourly. Which of the following considerations and configurations are relevant for optimizing cost and performance of the Cortex Search Service in this scenario?
- A. The
- B. For optimal performance and cost efficiency, Snowflake recommends using a dedicated warehouse of size no larger than MEDIUM for each Cortex Search Service.
- C. The primary cost driver for Cortex Search is the number of search queries executed against the service, with the volume of indexed data (GBImonth) having a minimal impact on overall billing.
- D. For embedding text, selecting a model like

- E. CHANGE_TRACKING
Answer: A,B,D,E
Explanation:
Option A is correct because a Cortex Search Service requires a virtual warehouse to refresh the service, which runs queries against base objects when they are initialized and refreshed, incurring compute costs. Option B is correct because the cost of embedding models varies. For example, 'snowflake-arctic-embed-m-v1.5 costs 0.03 credits per million tokens, while 'voyage-multilingual-2 costs 0.07 credits per million tokens. Choosing a more cost-effective model like 'snowflake-arctic-embed-m-v1.5 for English-only data can reduce token costs. Option C is correct because Snowflake recommends using a dedicated warehouse of size no larger than MEDIUM for each Cortex Search Service to achieve optimal performance. Option D is correct because change tracking is required for the Cortex Search Service to be able to detect and process updates to the base table, enabling incremental refreshes that are more efficient than full re-indexing. Option E is incorrect because Cortex Search Services incur costs based on virtual warehouse compute for refreshes, cost per input token, and a charge of 6.3 Credits per GB/mo of indexed data. The volume of indexed data has a significant impact, not minimal.
NEW QUESTION # 138
......
Free GES-C01 Exam Dumps to Improve Exam Score: https://www.actualtestsquiz.com/GES-C01-test-torrent.html
Exam GES-C01: New Brain Dump Professional - ActualTestsQuiz: https://drive.google.com/open?id=1uT-1lH2DpE3SD45eGeuz7WaLIP5WMO3M

